[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


Somebody make a competent collage script already edition

Discussion and Development of Local Image, Video, and Music Models

Previous: >>109599534

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg
>>
gm saars
>>
>mfw Resource news

08/19/2026

>Minimax H3 Latent Upscaler
https://huggingface.co/LBH-123-AI/Minimax_h3_latent_Upscaler

>CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
https://coinve200k.github.io

>LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
https://github.com/QHR69/LinCa

>RADmesh: Remesh-Aware Mesh Deformation
https://threedle.github.io/radmesh

>MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
https://github.com/facebookresearch/moe_vie

>Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models
https://github.com/cheny02/ADAPT-ACMMM2026

>Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models
https://github.com/Zane-ZYQiu/entry-point-umm

>aDSL: Agentic 3D Creation via Joint Agent-Program Design
https://github.com/sig-pku/aDSL

>Live Interactive Training for Video Segmentation
https://youngxinyu1802.github.io/projects/LIT

>ComfyUi-MiniMax-H3-Image-And-Reference-To-Video: Use I2V and reference images simultaneously
https://github.com/BigStationW/ComfyUi-MiniMax-H3-Image-And-Reference-To-Video

>ComfyUI-MiniMaxH3Mod: RefMod reference adapters
https://github.com/Luisacaotica/ComfyUI-MiniMaxH3Mod

08/18/2026

>Qwen-Video-Edit: Instruction-Based Video Editing by Repurposing an Image Editing Model
https://yunpeng1998.github.io/Qwen-Video-Edit-Page

>ByteDance just released Bernini‑Diffusers‑v2
https://huggingface.co/ByteDance/Bernini-Diffusers-v2

>PixRestore: Unified Image Restoration via Pixel Diffusion Transformer
https://github.com/csslc/PixRestore

>PixelControl: Fine-Grained Condition Fidelity in Text-to-Image Diffusion
https://linxin0.github.io/pixelcontrol_homepage/pixelcontrol-site

>MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling
https://expmaster.github.io/megaparts_webpage
>>
>mfw Research news

08/19/2026

>AI models can't tell time or read a calendar
https://www.livescience.com/technology/artificial-intelligence/ai-models-cant-tell-time-or-read-a-calendar-study-reveals

>From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
https://arxiv.org/abs/2608.18076

>MSEditor: Toward Consistent Multi-Shot Video Editing
https://arxiv.org/abs/2608.17559

>Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization
https://arxiv.org/abs/2608.18040

>SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
https://arxiv.org/abs/2608.17426

>AViTS: Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation
https://arxiv.org/abs/2608.17995

>TINA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Consistent Text-Free Inversion
https://qianlong0502.github.io/TINA-Plus-Homepage

>Magnitude-Direction Decoupling for Fast Video Generation with Flow Matching Models
https://arxiv.org/abs/2608.17695

>SE-MoLoRA: Shared-Expert LoRA Adapters for Domain-Specific Photographic Assessment
https://arxiv.org/abs/2608.17514

>GenRec: Knowing Where to Reconstruct and Where to Generate
https://arxiv.org/abs/2608.17832

>EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing
https://arxiv.org/abs/2608.18063

>SFMformer: A Spatial-Frequency Modulation Transformer for Lightweight Image Super-Resolution
https://arxiv.org/abs/2608.17966

>MS-MFAD : Multimodal large language models for Face Anti-spoofing Detection
https://arxiv.org/abs/2608.17328

>What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
https://what-to-edit-next.github.io
>>
whats the best turbo lora so far?
>>
>>109601850
maybe you do have a GPU on your CPU or its some misdetection? don't expect all of this python-torch-cuda-driver etc stuff to work properly in detail it always has been a mess
>>
>>109601864
>>109601887
>>109601893
I don't listen to avatar obsessed troons
>>
oh la la...
>>
The ravages of HRT.
>>
>>109602044
indeed
>>
Have you tried ltx 2.5? Minimax is so slow.
>>
hello
>>
>>109602268
hi
>>
>>109601828
i hate these low fps op images. you should delete this thread and make a new one without them. you should never make a new choppy fps op pic
>>
What does attention kitchen even do?
>>
>>109602342
Uses fast kernels.
>>
can somebody repost the miku image at the bottom from the op's imagine?
>>
>>109602374
It's still up in the last thread.
>>
File: h3_00117.webm (1.81 MB, 1808x1016)
1.81 MB
1.81 MB WEBM
>>
File: 1765603063909509.jpg (173 KB, 929x720)
173 KB JPG
kek, used a concord image for this one.

https://files.catbox.moe/isaql0.mp4
>>
>>109602397
lmfao
now generate her giving birth to quintuplets
>>
Janny didnt read and inspect the thread again. Just make a new one
>>
>>109602397
https://files.catbox.moe/c62we4.mp4
>>
>>109601828
https://n.uguu.se/pTNbdgDq.mp4
>>
>bought a brand new expensive AMD gpu with 32gb vram
>delivered yesterday and tried it
>assumed it needed a regular 12pin instead of a 12V-2x6
>only ran on like 15w
>motherboard was struggling to pick up its existence
>had to manually type in some linux code babble to make it be properly detected
>had to manually write more linux code babble for llama.cpp to pick it up
>fiddled with it for like 6 hours before figuring out I need another cable
holy shit, just stick to nvidia man, never fall for the AMD meme, the new cable should arrive today and hopefully it will fix everything
>>
>>109602493
we told you this so many times anon. you didn't listen
>>
File: goblin_mode.jpg (260 KB, 1920x1080)
260 KB JPG
>>109601828
goblin mode
https://youtu.be/7wOO9LNNNuAhttps://suno.com/s/d5JPwNPYoQukSqjX
>>
totally normal :)
>>
DUMB JAN
DO IT AGAIN
>>
>>109601828
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
File: 128518369816.jpg (315 KB, 1262x1246)
315 KB JPG
>generate a girl
>call it a boy

https://files.catbox.moe/gg8gto.mp4
>>
File: daboogie.gif (2.23 MB, 498x336)
2.23 MB GIF
>>
https://files.catbox.moe/oz8k0h.webm
>>
>>109602484
pretty common to have some portraits there like of your ancestor (e.g. grandma) or other inspiration / memory / aspirational person. especially if there are pictures on the left or right.
of course you could also have some nipponese calligraphy or a motto or some woodblock art reprint.
>>
>>109602569
One has a collage and the other has some guy who keeps spamming his retarded gorilla video. Also this one came first.
>>
Running this on minimax h3. I don't know what the fuck I'm doing and can't write NSFW prompts

https://files.catbox.moe/0gmr26.mp4
>>
you should all be ashamed
>>
>>109602585
seems to work except for you messing up the aspect ratio somewhere?
>>
File: kekekekkeeeeeeeeek.png (1.17 MB, 864x1184)
1.17 MB PNG
nah JSID already bruh keeeeeeeeeeeeeeeek
>>
Dumb janny
>>
>>109602594
Yeah it doesn't automatically convert it

https://files.catbox.moe/xa1giz.mp4
>>
>>109602493
To be fair, not researching what cable you need is on you.

Hope you're just trying to run LLMs with this thing, because image or video diffusion performance will be pretty lackluster.
>>
whats the best turbo lora right now?
>>
>>109602622
Just sick with lightx2v. I think it's the only ref one anyway.
>>
Someone please tell me where to get NSFW prompts. I don't know how to write.

https://files.catbox.moe/yblgah.mp4
>>
File: file.png (163 KB, 1117x833)
163 KB PNG
>>109602585
>>109602601
>>
>>109602624
which repo is it in?
>>
File: 154326.jpg (120 KB, 966x504)
120 KB JPG
>>109602585
You probably fucked up first frame resoultion and output resolution
>can't write NSFW prompts
Don't write it literally. The model can do nsfw but you must describe the details of the action, a single keyword is not going to cut it
Also it can't draw vagine and benis well so.
>>
N
>>
File: 1763169279121826.png (1.49 MB, 960x1280)
1.49 MB PNG
>>
>>109602590
You're heck'in cute and valid
>>
>>109602601
you can make it do that on the input reference(s) or generated video

probably better than getting very squished video? at least I think few people like those, it also sort-of affects what details you can see at some given resolution. you do you of course.
>>
File: 1783951348168839.png (1.53 MB, 960x1280)
1.53 MB PNG
>>
>>109602637
Google or hf is but a click away
>>
>>109602643
if you arent gonna post mimes then please leave
>>
>>109602646
ya think?
>>
File: 1785275927407853.png (1.48 MB, 960x1280)
1.48 MB PNG
>>
File: 1755783169109435.png (1.23 MB, 960x1280)
1.23 MB PNG
>>
bake without the rentries from now on. we can finally be free from wanschizo and ranfaggot slopgens
>>
File: 1757898576122926.png (1.36 MB, 960x1280)
1.36 MB PNG
>>
File: 1786931647910838.png (1.5 MB, 960x1280)
1.5 MB PNG
>>
File: 1759249186482383.png (1.52 MB, 960x1280)
1.52 MB PNG
>>
File: 1784167737818511.png (1.31 MB, 960x1280)
1.31 MB PNG
>>
>>109602629
afaik H3 isn't for lewds, you can use img ref tho
>>
File: gaddafi.jpg (8 KB, 277x182)
8 KB JPG
>the jeet has even more bots under his command now or local jeets he hired to shitpost
>reposts old shit without promoting his cause
>>
>>109602680
what can i use to make NSFW? Anime or realistic
>>
>>109602683
there is a coordinated effort to destroy /ldg/ collage threads because someone is seething that his gens dont get into it
>>
>>109602683
>reposts old shit
that thread got pruned
>>
>>109602654
make me
>>
Test 1 2 3.
>>
new loras for fl2v! just looked at the dir, 1.1 is up.

https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_fl2v_turbo_4step_v1.1_768p_comfyui_bf16.safetensors
>>
>>109602680
Lewd stuff works great, it's porn where it gets a little trickier.
>>
The usual collage autist supports keeping the rentry links in the OP.
Ani is tactically trying to create a split between the collage autist and the people who don't give a crap about always having a collage. His sole motivation is removing the rentry links, and will do whatever it takes to persuade people to want them removed.

It's not going to stop me from baking threads with the rentry links firmly in place.
>>
>>109602696
>bf16
too big to run
>>
>>109602665
You're still not a thread authority "anon"
>>
WHAT THE FUG!
I SAID PULL DOWN YOUR UNDERWEAR

I think my minimax is censored and cucked.

https://files.catbox.moe/6jnr5t.mp4
>>
>>109602696
wtf is going on with huggingface download speeds today, so slow
>>
>>109602702
no but I like seeing the self imposed hall monitors get dunked on by jannies
>>
File: 1783051870959260.jpg (495 KB, 832x1216)
495 KB JPG
>>
>>109602684
wan+older ltx tunes are more developed for full porn

h3 can make nsfw but even with lora a lot of it is currently coarse, lots of descriptions and not quite performing ideally probably meaning quite a lot of re-gens to get more ideal results.
>>
>>109602703
you probably don't want to see what's underneath desu
>>
>>109602706
Yeah the xmas melty with hundreds of deleted posts of him was funny af
Imagine shitting yourself for weeks publicly
>>
>>109602563
please leave this thread alone tran
you have your own
>>
>>109602696
larry's is better anyway
light2x always has that slomo shit
>>
>>109602705
>today
i'm thinking mainly the LLM releases and apparently more and more people using them burdened HF for a while now?
>>
>>109602699
>>
File: 1770418734567670.jpg (481 KB, 832x1216)
481 KB JPG
>>
>>109602724
* i have fast dedicated fiber internet, it has been going on for a while now. they clearly DO move a lot of data if the download stats aren't fake tho.
>>
>>109602703
"pull down underwear TO EXPOSE her benis"
>>
File: 1759157145326430.jpg (254 KB, 768x1344)
254 KB JPG
>>
File: 1764948966053117.jpg (284 KB, 896x1160)
284 KB JPG
>>
File: 1761999376483819.jpg (276 KB, 800x600)
276 KB JPG
>>
File: 1766867823537653.jpg (256 KB, 768x1344)
256 KB JPG
>>
File: 1773447614572885.jpg (141 KB, 512x512)
141 KB JPG
>>
File: 1755940173091370.jpg (796 KB, 1024x1360)
796 KB JPG
>>
File: 1770316776379191.jpg (432 KB, 1024x1360)
432 KB JPG
>>
>>109602127
Bump. Vae decoding is 1/3th of total gen time on that garbage.
>>
File: 1777075751038595.jpg (1.26 MB, 1664x2432)
1.26 MB JPG
>>
File: g.jpg (402 KB, 1600x2400)
402 KB JPG
>>109602764
redoing some prompts on newer models?
>>
File: 1780582262227497.jpg (1.69 MB, 1664x2432)
1.69 MB JPG
>>
File: MiniMax-H3-00092.mp4 (3.54 MB, 864x480)
3.54 MB
3.54 MB MP4
>>
>>109602780
what do you want to know?
>>
File: 1756950972278483.jpg (465 KB, 1280x1024)
465 KB JPG
>>
>>109602796
> Have you tried ltx 2.5?
>>
>>109601828
i think everyone agreed we should add this to op

https://rentry.org/ranfaggot
>>
>>109602802
no ani, it's you and debo. That's the way it will always be.
>>
>>109602801
yes, i have
>>
>>109602802
Yep. We should get rid of the schizo that keeps defaming Ani. I agree.
>>
>>109602802
agreed
>>109602805
kys catjack
>>
>>109602802
that's a good idea, we should also add anistudio when ani says it's ready, that'll make the schizo seethe
>>
>>109602805
everyone knows you are catjack retard
>>
>>109602802
no, I fight against any rentry
>>
>>109602806
Is it any good? Does it support image, audio and video references like minimax? Is it fast?
>>
>>109602817
>everyone
In your head?
>>
>>109602667
this is awesome, can i get a catbox?
>>
File: 1782696906953517.jpg (1.37 MB, 2560x1440)
1.37 MB JPG
>>
File: 1758210459842890.jpg (700 KB, 1920x1088)
700 KB JPG
>>
After all these decades there's still no defense against someone batching uploads to spamfill a thread to death?
>>
File: 1779207351455173.jpg (398 KB, 1152x1152)
398 KB JPG
>>
>>109602830
https://files.catbox.moe/1wr3ss.png
>>
File: comparison.webm (3.5 MB, 1300x960)
3.5 MB
3.5 MB WEBM
>>109602696
4 steps 1.1 lora (left)
32 steps no lora (right)
>>
File: 1767762329029052.jpg (300 KB, 832x1216)
300 KB JPG
>>
>>109602825
>Is it any good?
it's better
>Does it support image, audio and video references like minimax?
kind of. minimax was specifically trained for it while ltx requires you to condition the context window. regardless, you can insert images, audio, or even a video at the start of the context window and then prompt your real scene. if you used the fl2v h3 model, then it's exactly like that.
>Is it fast?
i can generate 5 or 6 videos of the same length and resolution with better image quality by the time h3 can finish one of them
>>
>>109602843
Where does he get the images to spam?
>>
File: 1774490208557877.jpg (1.49 MB, 2704x1544)
1.49 MB JPG
>>
File: 1785410626622459.jpg (838 KB, 1248x1824)
838 KB JPG
>>
>>109602863
extend...
>>
File: 1766643408194865.jpg (710 KB, 1248x1824)
710 KB JPG
>>
>>109602854
I'll give it a try then. Why does no one post ltx 2.5 gens here?
>>
>>109602871
I'm a vramlet and it takes longer
>>
File: 1761781244761802.jpg (720 KB, 1248x1824)
720 KB JPG
>>
File: 1778843617295048.jpg (326 KB, 888x1184)
326 KB JPG
>>
>>109602871
>Why does no one post ltx 2.5 gens here?
i don't know. the goyim can not be helped
>>
>>109602863
get fuckin banned bitch
>>
>mfw Research news

>SafeCA: Safe Cross-Attention Localization and Regulation for Text-to-Video Jailbreak Defense
https://arxiv.org/abs/2608.10933

>TGRHuman: Text-Guided Realistic 3D Human Generation via Diffusion Renderer
https://arxiv.org/abs/2608.12175

>VidForensics-M1: Meta-Detection Reinforcement Learning with Verifiable Temporal Grounding for AI-Generated Video Forensics
https://arxiv.org/abs/2608.11201

>A Content-Aware Pure Permutation with Intrinsic Avalanche Effect: Breaking the Diffusion-Permutation Dichotomy
https://arxiv.org/abs/2608.09452

>Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite Avatars
https://arxiv.org/abs/2608.12107

>RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing
https://arxiv.org/abs/2608.09186

>CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
https://arxiv.org/abs/2608.13226

>SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data
https://arxiv.org/abs/2608.12876

>DiffSafeMerge: Mitigating Backdoor Inheritance in Diffusion Model Merging
https://arxiv.org/abs/2608.09445

>When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
https://arxiv.org/abs/2608.10489

>Preserve More Details: Mitigating Content Drift in Real-World Image Super-Resolution
https://arxiv.org/abs/2608.09373

>Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning
https://arxiv.org/abs/2608.09682

>Test-Time Hallucination Control in Large Vision-Language Models
https://arxiv.org/abs/2608.11474

>LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering
https://arxiv.org/abs/2608.07863

>Signpost Watermarking: Joint Optimization for Visual Watermark Coexistence
https://arxiv.org/abs/2608.10091
>>
>>109602871
Because it's just not as good as Minimax. It's a lot faster, and the resolutions are higher, but it can't handle complex scenes a lot better than LTX-2.3.
>>
File: 1762307544118871.jpg (173 KB, 640x448)
173 KB JPG
>>
>>109602881
the real news is that old reddit doesn't open anymore without an account
>>
File: 1782429042058791.jpg (172 KB, 512x512)
172 KB JPG
>>
File: 1758785540232718.jpg (277 KB, 576x576)
277 KB JPG
>>
>>109602900
who cares?
>>
File: 1768359112976156.jpg (289 KB, 1024x1408)
289 KB JPG
>>
>>109602909
I do
>>
File: 1781814228748183.jpg (70 KB, 384x640)
70 KB JPG
>>
File: 1782411575331079.jpg (303 KB, 576x576)
303 KB JPG
>>
File: vj88ffmpl24d1.png (99 KB, 475x465)
99 KB PNG
>>109602911
>I do
>>
>>109602920
> without an account
retard
>>
File: 1759715552537947.jpg (1.88 MB, 2048x3072)
1.88 MB JPG
>>
File: 1763405438562443.png (1.3 MB, 960x1280)
1.3 MB PNG
>>
File: 1786967801786514.jpg (307 KB, 704x1280)
307 KB JPG
>>
File: 1766451766878876.jpg (125 KB, 512x512)
125 KB JPG
>>
>>109602924
>retard
says the redditor
>>
>mfw MORE Research news

>FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait Image
https://kim-youwang.github.io/FiCA

>Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning
https://arxiv.org/abs/2606.24548

>TooBad: Backdoor Diffusion Models with Ultra-Low Poison Rate and Imperceptible Trigger
https://arxiv.org/abs/2606.23362

>DivRL: Disentangled Self-Similarity Rewards for Diverse Subject-Driven Generation
https://qianwangx.github.io/DivRL

>Semantic Browsing: Controllable Diversity for Image Generation
https://saradorfman1.github.io/SemanticBrowsing-webpage

>Geometry-Instructed Video Editing
https://geometry-instructed-video-editing.github.io/give

>The Professor: Multi-Teacher Unsupervised Prompt Distillation for Vision-Language Models
https://arxiv.org/abs/2606.23897

>UniRED: Unified RGB-D Video Frame Interpolation with Event Guidance
https://arxiv.org/abs/2606.24282

>Prompting Diffusion Models for Zero-Shot Instance Segmentation
https://arxiv.org/abs/2606.22660

>Trustworthy Image Authentication using Forensic Knowledge Graphs
https://arxiv.org/abs/2606.23917

>RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
https://arxiv.org/abs/2606.23221

>TuringViT: Making SOTA Vision Transformers Accessible to All
https://arxiv.org/abs/2606.24253

>Listening makes Vision Clear for VLMs
https://arxiv.org/abs/2606.23763

>Unmasking LAION-5B: Age, Gender, Race, and Emotion Biases in Large-Scale Image Datasets
https://arxiv.org/abs/2606.23204

>Fursee: Hybrid YOLO-DINOv3 Framework for Fursuit Identity Retrieval and Clustering
https://arxiv.org/abs/2606.22872
>>
>>109602897
so LTX is better if I just want some simple stuff, like character doing some action?
>>
>nigbo
>>
thread is full of anime 1girls and schizos arguing
well done guys you've scared away all effortposters with your bullshit now there's no one left here, i'm leaving too
>>
I'm about to setup a WAN 2.2 NSFW and generate porn with it.
I will post results when I get it running.
>>
>>109602985
>now
That point was actually like 3 years ago newfaggot
>>
>>109602985
Come back when it's American hours. Third worlders and Eurocucks makes bad company. They are poor and insufferable.
>>
File: crackhead.gif (1.53 MB, 498x498)
1.53 MB GIF
yall got any of them kinos?
>>
>>109602947
If you don't want any complex action or camera movement, that could work for you. Be mindful that it tends to shift 2d stuff into looking 3d.
>>
File: MiniMax_00002_s.mp4 (546 KB, 768x1376)
546 KB
546 KB MP4
>>
i am creatively bankrupt
>>
>>109603120
Find a sound clip from a favorite game or movie and use it as driving audio for a video. See where it takes you
>>
File: h3_00120.webm (3.24 MB, 1808x1016)
3.24 MB
3.24 MB WEBM
>>109603146
https://files.catbox.moe/p68yul.webm
>>
>>109601904
>>109602802
>>109602805
>>109602810
>>109602813
Hang yourself
>>
>>109603273
https://www.youtube.com/watch?v=PYKDyXnJl5E
>>
>>109603290
Localsisters. How will we recover?
>>
File: x.mp4 (1.47 MB, 480x864)
1.47 MB
1.47 MB MP4
>>109603120
still beating the average. in hollywood too.
>>
https://files.catbox.moe/j3gbrs.webm
>>
>>109603120
-browse the internet
-go watch a movie and rip off ideas from it
-same thing when you read manga or watch anime
-play video games too for ideas

skies the limit for ideas, just don't limit yourself and use this amazing AI model that China made and go nuts
>>
First one with WAN 2.2. It loops instead of going fully lewd

https://files.catbox.moe/egifs6.mp4
>>
>>109603394
why is it so damn fast
>>
why can these god damn namefags and avatarniggers not just leave us alone
>>
>>109603303
>>109603340
I'd play these.
>>
>>109603397
It's supposed to be a fast model. I was hoping it meant generates fast
>WAN 2.2 Enhanced NSFW | SVI | camera prompt adherence (Lightning Edition) I2V and T2V fp8 GGUF
>>
>>109603394
>First one with WAN 2.2
Bruh Wan 2.2 is over a year old at this point.
>>
>>109603070
I only do anime, so I'll stick to H3 then.
>>
File: x.mp4 (1.65 MB, 480x864)
1.65 MB
1.65 MB MP4
forgot to kill the the robowolf

>>109603405
soon
>>
File: Krea2_turbo_00888_.jpg (1.93 MB, 1672x2512)
1.93 MB JPG
Feels good to be fully rested and not pulling all nighters
>>
What does "MiniMax H3 Chunk FeedForward" do ????
>>
File: IMG_20260819_184959_504.jpg (507 KB, 2448x2448)
507 KB JPG
I turned a double-sided ID4 gen into a keychain.
A colorful chibi thin line art illustration of a cute very chubby hatsune miku, small breasts, wide hips, fat thighs, chubby ankles, fupa, wearing a frilly puffy black and white maid uniform, shy bashful pose with body turned to the side, legs crossed, hands clasped in front, head tilted down but eyes focused on the viewer, blushing, shy, against a white background. Divide the image into two equal left and right panels and create front and back views of the image, making sure the pose is mirrored accurately so the image can be printed double-sided.

Genning a 16:9 2.0 MP gives you two views you can then do some minor cleanup, checking, and flipping in something like Gimp. About 1/5 the time ID4 gives you a perfect alignment, you do have to roll the dice on it a bit.
>>
>>109603523
Can't think of more boring character than Miku.
>>
File: x.mp4 (1.58 MB, 480x864)
1.58 MB
1.58 MB MP4
>>109603533
> MiniMax H3 Chunk FeedForward — MLP peak-memory reduction, model_patches/memory
> Splits H3's token-local feed-forward over the packed sequence dimension while preserving ComfyUI's linear_input_act implementation inside each chunk. Independent of Sol-Attn — it works with any attention backend. Most effective on the INT8 ConvRot checkpoint, where the swiglu activation is fused into the INT8 quantizer and the [tokens, 28672] first projection dominates peak MLP memory; on a plain bf16 checkpoint the eager swiglu path holds extra intermediates and the saving is small.
>>
>>109603536
yup
>>
>>109603540
In english, doc
>>
>>109603545
memory reduction
>The benefit scales linearly with the packed sequence: the intermediate is 56 KB per token, so the saving grows from ~238 MiB at 8K tokens to ~1.9 GiB at 65K (measured, ×2 chunks, int8 checkpoint). Below min_tokens the node does nothing at all — lower it if you want relief on short renders too, raise it if you only care about long ones.
>>
>>109602802
Why not add the guy that keeps talking about these fucking people like we need a constant lore refresher when nobody else cares?
>>
>>109603547
Any quality reduction ? Without this node my gpu goes quiet and does nothing at middle of gen (on longer duration at higher res)
>>
Trying H3 local on 16GB RTX4080 and 32GB ram. I put 96GB paging on my nm790 just in case. But while ram is maxed out during runs, there's no drop in GPU utilisation and no spike in nvme access, doesn't seem to be any paging stall, so I don't think 64GB will help? I haven't tried going beyond 10s or higher res yet so maybe that might do it
>>
>>109603553
Haven't verified it but I think this one was lossless. AFAIK just some mild computational overhead for RAM savings.
>>
>>109603536
OK yes but she is useful because most models know her by default. Krea2 can do Migus well without a lora or a tedious description of her.
>>
>>109602802
you can tell who is the actual schizo retard when you compare the writings in that rentry vs the ones in the op, lol

also if we're talking about polls, there was a poll created by me (guy who doesnt care about any of the schizos anyway, but was interested to see a live vote of the people in the thread at the time about rentries), which got a quick 5 votes to keep the rentries, and showed that there were only 2 IPs in the thread that wanted the rentries removed while everyone else wanted them to stay. you can check the archive to see the voting timeline.

https://strawpoll.com/GeZARemG8yV
>>
File: Krea2_turbo_01215_.jpg (1.88 MB, 2368x1776)
1.88 MB JPG
>>109603573
He's just complaining to complain
>>109603577
It's sad because he keeps posting that while everyone that reads it calls him a bitter schizo. Damn shame he can't figure it out
>>
>>109602793
KWAB
>>
File: Poll.jpg (17 KB, 630x129)
17 KB JPG
Important data mining Poll

https://strawpoll.com/GPgVYm0LAna
>>
>pollschizo is back
>>
>>109603562
I think you are trying to gen something your system can't handle in reasonable time frame.
I have similar setup 4060TI 16Gb and 64Gb ram, I just gen less than 10s and 1000x1000 res vids
>>
File: migu-kawaii.mp4 (1.77 MB, 864x480)
1.77 MB
1.77 MB MP4
>>109603607
Well yeah some people get "triggered" by Migus (or anything the don't like) which is another reason to post them.
>>
>>109603644
>half the models missing
>>
>>109603659
People need help, they also tend to not add anything other than negativity, not even a gen.
>>
File: minimax_00026-audio.webm (1.85 MB, 928x672)
1.85 MB
1.85 MB WEBM
>>
>>109603647
I do not know the lore of this board, I'm merely visiting to drop this poll
>>109603661
It's about picture gen
>>
>>109603644
i still use chroma
>>
>>109603649
What models and settings? I'm just trying basic stuff at the moment using image 2 video and ref 2 video on the H3 MiniMax int8 models in comfyui
>>
>>109603675
>It's about picture gen
yes, and half the models are missing
>>
>another day of worsening my ED with H3
feels soft, man
>>
>>109603672
This time it actually looks like a proper 1:1 scene from the game.
>>
>>109603672
Ahahah! Yes! I knew you'd have her run over Teto at the end!
>>
File: 1786501250536323.jpg (37 KB, 335x597)
37 KB JPG
>Error: You were warned. You must first view this warning to post again.
>>
>>109603677
I use int8.
>>
>>109603710
Don't badmouth the raped retard or it'll sic its army of proxies to report your post, anon
>>
3 gens in a row without body horror with H3, that's like winning the lottery
>>
>>109603740
Only time I've seen body horror with H3 is if I've repeated a body part too much so I get extra feet
>>
>>109603740
>t. speedcope lora user
>>
>>109603740
It's the 1/500 legitimate horror gens that get me.
Like broken audio, strange artifacting, etc. Straight out of the ring.
Creeps me out enough to stop genning.
>>
>>109603756
I only use sage, it changes the gen yes but only just barely
>>
>>109603677
>>109603715

10 steps and euler, other settings I keep at default.
>>
>>109603759
>Creeps me out enough to stop genning.
weakling. you haven't seen what i have seen
>>
>>109603770
Oh yeah... I remember trying to do porn in wan and getting stuff like the subject pulls the skin off their leg instead. T... thanks... wan team.
>>
H3 is great but doesn't feel quite there yet for more serious longer projects. Can't wait to see what they do with the next version.
>>
File: Krea2_turbo_01242_.jpg (1.89 MB, 2368x1776)
1.89 MB JPG
>>
>>109603783
did you see the flesh?
>>
>>109603691
https://files.catbox.moe/g4p5kh.mp4
>>
>>109603798
>soft r
>>
>>109603740
So glad I'm not trying to wring porn out of H3. Never had issues with body horror, only the occasional additional arms in some gen referencing a video.

>>109603784
You'd need some really good references (like 360 shots of every location) and meticulous planning for that. Absolutely takes the fun out of genning.
>>
>>109603770
Man, I wish I hadn't deleted the most recent one. I was doing standard porn gens, then out of nowhere got a complete distorted zoom in of the characters face and upper body, all blown out and distorted, but the eyes were strange and scan-liney. The entire video has 15 seconds of them talking nonsense in distorted slowmo while things crawled out of their stomach and dropped out of view. Definitely made my skin crawl.
>>
>>109603805
https://files.catbox.moe/06gilf.mp4
>>
>>109603811
kino
>>
>>109603795
yeah like horror movie gore, I'm sure it was deliberate to fuck NSFW
>>
>>109603807
>Absolutely takes the fun out of genning.
I've made 3 low-effort music videos with AI now and I wouldn't say it takes the fun out of genning because you're in an interesting middle stage after the boring beginning of collecting references and prototyping

The actual worst part is editing it all together at the end because the 80% that takes 20% of the time feels awesome and like you're making fast progress and then the last 20% that takes 80% of the time shows up and you objectively do not have fun while going through the motions and grinding out the rest
>>
File: x.mp4 (1.94 MB, 480x864)
1.94 MB
1.94 MB MP4
>>109603784
I do think people already will make "serious longer projects" with it, whatever that actully means.
>>
File: x.mp4 (2.42 MB, 608x1056)
2.42 MB
2.42 MB MP4
one for you >>109603405
>>
>>109603862
>whatever that actully means
For me that's full episodes, short films, and feature-length films. The first 2 are maybe possible but the process and result will be an inconsistent mess. There needs to be improvements at both the model level and frontend. A dedicated video gen frontend instead of comfy is the way imo.
>>
>>109603672
Based
Teto is shit as a waifu anyway
>>
>>109603791
box pls?
>>
File: test_00045_.png (1.27 MB, 1536x1536)
1.27 MB PNG
>>
i been away for a while. anything new for local gen goons? i have a 5070ti and 64gb of vram, last time I genned I'd use wan2.2
>>
File: mikan-rito_00016_.png (1.27 MB, 1024x1536)
1.27 MB PNG
>>
Best H3 model or lora for realistic lewd stuff? Base model is surprisingly good at boobs but not anything else. pinkcherry and 10Eros also seem to fail at anything beyond boobs.
>>
>>109603837
I only did one music video, but it was indeed quite fun, because some dancing or basic movement is easy to get right and you don't need any real continuity. Barely did any editing except for throwing the videos together and timing the music somewhat well. Will probably do more, even though people don't really care about AI slop music videos.

https://files.catbox.moe/hnbyt6.mp4
>>
>>109603944
Yes, download Minimax H3 right now.
It's 500 times better than wan2.2 for generating video.
>>
>>109603886
give the people that already made 5+ minute long movies with wan free access to better hardware and they might just do a 1.5h film, heh
>>
File: imgSP_00101_.jpg (717 KB, 1376x2008)
717 KB JPG
>>
Anon who wants to create an anime here.

So I'm starting to craft my references for environment, BGM, voices, characters, etc.
What is the correct way to ensure a height difference in H3? And maintain character height consistency? With the reference sheets for characters they just fill up the height of the output image so I need some way to make sure certain characters are a fixed value taller or shorter than another character.
>>
>>109603960
catbox?
>>
i WOULD make a long kino... but i'm not gonna stop tweaking this little prompt until i get it right
>>
File: x.mp4 (788 KB, 480x864)
788 KB
788 KB MP4
>>109603986
probably some puppet/skeletal rig video reference?
>>
>>109603986
Could help adding something common of known height to their reference sheet, like the universal tool for identifying manlets in photos: a door with a visible door handle
>>
>>109603960
image models are now fucking great
>>
File: 1778959937455165.jpg (62 KB, 620x624)
62 KB JPG
Comfy UI v0.27.0
>Prompt executed in 254.57 seconds

Comfy UI v0.33.0
>Prompt executed in 310.54 seconds

WHAT THE FUCK MAN ???????
>>
>>109604072
>>109604079
>>109604091
>this obvious samefagging
>>
File: Krea2_turbo_01290_.jpg (1.62 MB, 1776x2368)
1.62 MB JPG
So desperate and nothing of value created, what's sad is 80% of these events happen while I'm on hiatus
>>
File: 1774470289462675.png (1.07 MB, 960x1280)
1.07 MB PNG
>>
File: 1759766801719277.png (1.31 MB, 960x1280)
1.31 MB PNG
>>
reminder that debo won since you are posting in his thread
>>
File: imgSP_00118_.jpg (1.07 MB, 1512x2208)
1.07 MB JPG
>>109604057
agreed
>>
File: 1774397613744695.png (859 KB, 960x1280)
859 KB PNG
>>
File: output_no_audio_lossless.mp4 (3.36 MB, 1056x608)
3.36 MB
3.36 MB MP4
>>109604130
So pulling all nighters to hijack a OP for weeks on end is a win?
Really sad when you think about it
>>
>>109604146
>pulling all nighters
you were there too
>>
>>109604146
why are you are such a drama-obsessed loser? is it some ghetto thing?
>>
What a shitty low VRAM thread.
>>
File: Krea2_turbo_01308_.jpg (1.67 MB, 1776x2368)
1.67 MB JPG
This is satire at this point, not engaging
>>
>>109604167
corr
>>
>>109604167
It's so tiresome...
>>
Imagine making a lora and thinking it's not checksum hashed and uploaded to a database by the Lora node. lmao
>>
>>109604167
nice fused leggings fucking retarded spammer
>>
We should have a diffusion convention and invite debo.
>>
>>109604215
>>
thanks anon^
>>
>>109604226
why so fucking early? and troll links again?
>>
>>109604130
Boi won so hard he gets insulted in every timezone on sight
>>
someone bake the real thread
>>
>speed running an early bake
>>
tfw the model knows the office really well but not parks and rec
>>
File: 1771834961334218.jpg (85 KB, 512x512)
85 KB JPG
>>
should have baked faster lol
>>
>>109604226
An IRC mod confirmed to me they "usually go with whichever thread was made first".
So that sets the precedent. Nothing wrong with early threads if it means sticking it up that dumbfuck schizo ani.
>>
File: gemma_lmg_lick.mp4 (628 KB, 1080x620)
628 KB
628 KB MP4
So, does /ldg/ have a mascot like /lmg/?
>>
>>109604274
bro is pretending to talk to the mods kekekek
>>
File: 1761996893226331.jpg (102 KB, 512x640)
102 KB JPG
>>
File: imgSP_00124_.jpg (783 KB, 1384x2064)
783 KB JPG
>>
File: 1756026062506400.jpg (98 KB, 512x640)
98 KB JPG
>>
>>109604283
I literally spoke with an IRC mod called Faust just a few hours ago, and he told me exactly that. Also said the other thread may have been deleted by mistake, and rejected the janitor's ban request for my Spamming/Flooding ban.
This isn't rocket science you moron, the IRC has been around for a long time, and the mods there have had to become more receptive to inquiries since the feedback form was shut down.
>>
File: 1766364637267875.jpg (623 KB, 2048x2048)
623 KB JPG
>>
File: 1766798601143921.jpg (114 KB, 512x640)
114 KB JPG
>>
>>109604306
mods don't discuss things like this. you are mentally ill if you believe that you experienced this
>>
File: 1784032563325964.jpg (127 KB, 384x896)
127 KB JPG
>>109604306
be thank
>>
File: 1765883027040817.jpg (910 KB, 1536x1536)
910 KB JPG
>>109604306
mods don't
>>
File: 1785849411399063.jpg (93 KB, 512x512)
93 KB JPG
>>109604274
bro is pretending
>>
>>109603986
I would prepare a sheet with the rest of the characters against the protagonist as another reference. Or if two characters are going to be in a scene, I'd put them side by side as a reference. And no matter how much you give it H3 is bound to screw things up, luck plays a part.
>>
File: 1768718377176991.jpg (908 KB, 2048x2048)
908 KB JPG
>>
File: 1786991808470322.jpg (227 KB, 1024x1024)
227 KB JPG
>>
File: 1770000487063410.jpg (144 KB, 512x512)
144 KB JPG
>>
File: 1755980933486217.jpg (317 KB, 1079x1074)
317 KB JPG
>>
File: 1768587662439058.jpg (102 KB, 512x512)
102 KB JPG
>>
File: 1763058629463568.jpg (185 KB, 512x512)
185 KB JPG
>>
>>109604282
yes, its a fat black man
>>
>>109604306
Could tell them all this spam is due to a wothless tranny retard wanting to advertise its dogshit wrapper (which is against the rules) and trying to hijack threads for that purpose.
Maybe then they will do something .
>>
>>109604203
Ah the old "Hahaha! You made a spelling mistake! I WIN!" kind of reply. Keep up the good work.
>>
>>109604486
.
>>
>>109604509
?
>>
>>109604466
No point, they're not interested in the autism dynamics of generals across the site, and barely care to engage with users as it is. I can tell you for a fact that mods already spend a lot of time and effort telling janitors not to try and identify specific users/trolls and only focus on posts that break the rules; they have even less time for users trying to point this out over the IRC.
>>
>>109604525
bro is still larping KEKEKEK
>>
File: 1769209720508386.png (24 KB, 355x310)
24 KB PNG
>>109604537
Trani.......
https://4chan.org/faq#banappeal
>>
>>109604562
that isn't a chatlog
>>
>>109604588
It sure isn't, Julien, it sure isn't
>>
>>109604629
try again later
>>
File: 1782386655894734.jpg (58 KB, 801x89)
58 KB JPG
>>109604339
>mods don't discuss things like this. you are mentally ill if you believe that you experienced this
damn i guess trani is mentally ill then
>>
File: 1763032374047165.jpg (151 KB, 801x387)
151 KB JPG
>>109604679
>>
>>109604679
>>109604689
>All that effort to get a name only for it to be its own
what an incompetent retard
lmao
>>
>>109604679
i was referring to your imagination of mods discussing janitor decisions
>>
File: 26820.jpg (502 KB, 1712x1200)
502 KB JPG
>>109603955
love the idea, execution is a bit lacking imo hope you keep at it
>>
File: 426951390546426.mp4 (3.74 MB, 1120x480)
3.74 MB
3.74 MB MP4
>>
File: MiniMax_H3__00197.mp4 (3.09 MB, 576x896)
3.09 MB
3.09 MB MP4
>>
File: 1763543607054314.png (5 KB, 139x304)
5 KB PNG
>Hey let's try finally enabling dynamic VRAM on comfy, it has to be fixed by now for my 24GB VRAM 128GB RAM machine!



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.