Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109475019https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg
based
zased
>>109476286>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
>mfw Resource news08/05/2026>Inline Studio v1.2.62 - Minimax H3 Lora training still onlyhttps://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3https://huggingface.co/Kijai/MiniMax-H3-TAE>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inferencehttps://github.com/6somehow/DAC-SPADE>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generationhttps://github.com/yizzz927/CAPE-T2V>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusionhttps://github.com/jd-opensource/JoyAI-Video-Edit>ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMshttps://github.com/YangYangGirl/ParVL>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diethttps://huggingface.co/JamesZar/OliveGemma-3B08/04/2026>stable-diffusion.cpp adds support for MiniMax-H3https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES timehttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editinghttps://github.com/IntMeGroup/MIEScore>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videoshttps://rathgrith.github.io/PeCA>Kandinsky WM 1.0: A family of models for Physical AIhttps://github.com/kandinskylab/kandinsky-wm08/03/2026>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>109476286will smith is so stronk the jannies can't do anything against him!
>mfw Research news08/05/2026>SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrievalhttps://arxiv.org/abs/2608.03120>HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Planehttps://arxiv.org/abs/2608.03422>DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformershttps://arxiv.org/abs/2608.03082>Can T2I Models Draw from the Right Frame of Reference?https://arxiv.org/abs/2608.03357>Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editinghttps://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0>Self-Supervised Representation-Guided Generative Dataset Distillationhttps://arxiv.org/abs/2608.03218>Latent Reward Registers for Diffusion Preference Alignmenthttps://arxiv.org/abs/2608.03929>UniWorld-Design: From Pixel Generation to Layer-Native Designhttps://arxiv.org/abs/2608.03971>MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Bindinghttps://arxiv.org/abs/2608.03708>RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editinghttps://arxiv.org/abs/2608.03059>Efficient Video Dataset Distillation via Cluster-Guided Prototype Blendinghttps://arxiv.org/abs/2608.03269>Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfoldshttps://arxiv.org/abs/2608.03135>TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Modelshttps://arxiv.org/abs/2608.03057>Adaptive Two-Stage Visual Token Pruning for Efficient Inference in VLMshttps://arxiv.org/abs/2608.03112>Enhancing VLM Reward Models Through Structure-Aware Fine-Tuninghttps://arxiv.org/abs/2608.03875>Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understandinghttps://qwen-3d.github.io>When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardwarehttps://arxiv.org/abs/2608.03649
me rn
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lorathey say its not fully trained yet, but something to test, the full one will obviously be better
the nice thing about ltx is that all the little micro camera movements were implicit. with h3, i have to put timestamps all over the place telling it to look here and there
>>109476297interesting read. this "animanon" seems like a bit of a faggot.
and suddenly its as if all the "newfrens" have vanished quite the coincidence
>>109476331totally
>>109476346I'm still here
>>109476346vague king
>>109476353that is impressive
>>109476346im totally new here and im wondering why anon is spamming off topic links??? >>109476297 >>109476297 >>109476297how are those links related to local diffusion?????? seems to me like they just defame a literal saint and programming god who will destroy comfyui am i right??
>>109476357>replying to yourself
I was too lazy to test H3, maybe on weekend..
>>109476368You are missing out
>>109476368it's a nothing burger, don't bother
>>109476357>a literal saint and programming god
>>109476297>>109476331>>109476347Samefag.
>>109476357How much is tr(ani) paying you
>>109476368kek why? it's an amazing model and comfyui has templates, wan2gp has it preconfigured. well worth your time if you're ever entertained by such models I'd say
>>109476386>no u
>>109476286Can someone give me a good prompt for minimax H3 ref2v workflow. I just want to change the character in the video with the character in the reference image.
>>109476405You're not attempting to do anything untoward, are you?
>>109476405character swap is a shot in the dark
>>109476386proof?
>>109476346almost as if "someone" isnt trying to hit bump limit as quick as possible anymore huh
>>109476389idk, tired in general
>>109476405You want to use the Load Video -> Get Video Components nodes and align them to the R2V workflow. Then in the prompt, you want to specify that you want to replace <Subject 1> in <Photo 1> with whoever from <Video 1>. Need to test this in the morning to see if it actually works, but let me know anon.
>>109476356agreed. that was a simple prompt (draw graffiti, followed by pose)looks like we could probably even get her to change cans if i fully described each color she should pick up/set down? not that I will do that rn
>>109476444good trips, very good post, shit gen but it gets the point across just right
Is the 4-step lora only for the fl model? https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora
I understand the desire to go fast but is it truly worth quality loss to use a turbo lora? Sure I understand when its wait a minute or two vs half an hour, but even then surely its not needed?
>>109476467ignorance is bliss
>>109476467you use it if you are tweaking a prompt. the motions are mostly correct with the lora so you can crank it up when it seems right
kino, reference works well once you figure it out<Subject 1> is the character in <Picture 1>. A high-detail cinematic tracking shot of <Picture 1>. He is walking through a dimly lit office with a "Halo Studios" sign on the wall. A man with purple hair and a tshirt saying "they/them" walks up to <Picture 1> and says in a high-pitched, frantic, youthful voice "you didn't respect my pronouns!". The camera focuses on <Picture 1> as he speaks in a deep, gravelly voice: "I don't give a fuck". His tone, pace, and vocal style perfectly match <Audio 1>, with precise lip-syncing. <Picture 1> fires his halo battle rifle at the man with purple hair standing 5 feet away, and they fall to the ground. <Picture 1> walks out the doors of the office.https://files.catbox.moe/lveb7c.mp4
>>109476485that's pretty simple, two of the four basic headers
>>109476485audio tip: 5-10 seconds but not over 15 for cloning, then use that>>109476489true but you also need to specify voice traits for the other characters or they will both get the voice, for more complex stuff google ai mode already knows the minimax manual stuff.
>>109476485Any luck with using it as an image edit model?Currently trying to do a 0.5 second video where given a base image, the character is replaced with a reference image but no luck so far. It partially works but not as good as nanobanana would normally.
>>109476485oh you know what you could do? >>>/v/744650971this could be some real kino
Can you specify height of characters
>>109476444yeah stop motion looks much better
I know you guys are all into your minmax vids but I'm having issues with just generating a simple image. I haven't messed around with any of this stuff since A1111 in 2023. Is this a sampler issue, a steps issue, cfg? Is there some sort of pre-made workflow for realistic z-image gens?
>>10947649610/10 I'll take a dozen
>>109476435stretching out the internet shitposting is not da wae to get motivation, nap/sleep and do it after you wake up again
>>109476495<subject 2> a massive man twice the height of <subject 1>
anyone else been running some of their shit through starlight precise?
>>109476496New Robot Chicken ep just dropped
KINOhttps://files.catbox.moe/u060cd.mp4
>>109476500Use some official workflow. That weight dtype you set for the model may be the cause as well.
>>109476508probably not. because for what purpose even?
>do a final fantasy i2v from a game screenshot>it does the broken screen space reflections properly as the main subject movesholy fuck
this is actually a realistic scenario btw:https://files.catbox.moe/ya0wtt.mp4
>>109476517I did that but it puts everything into a single node and doesn't let me use loras. Also the weight-dtype was just me fucking around, I used default for everything previously.
>>109476529fuck
6 steps euler/beta with turbo lora. Audio seems fine.https://files.catbox.moe/e157fm.mp4
>>109476530>it puts everything into a single nodeenter the subgraph (it's a button on one of the top corners of the subgraph node, forgot if left or right), or unpack it (right click menu)
>>109476453
Anons, H3 has been out for literal days. What's the strategy to get it to generate actual sexy shit and not weird SD 1.0-style body-horror?
>>109476346No Im just lurking after spending a day trying to get my gens sharp. Trawling through other peoples workflows atm
>minimax_h3_video_vae_int8_convrot.safetensorsI get black screen on final video with this vae
>>109476569did you update comfyui retard-chama
>>109476297>https://rentry.org/animanonIs it true that this guy is one of H3 devs? I've heard that he flew to Japan to strike a deal or something. The name makes sense
>>109476569update comfy / CU130+ / comfykitchen
>>109476529It can do stop motion? Unedited?
>>109476519Better quality and less compute but its probably not even necessary if you're a cartoonfag.
>>109476574No, Ani, your name is because you used to make dogshit gifs in the SD 1 days. You have nothing to do with h3, absolutely zero model baking skills, and you flew to Japan to beg for money only to get laughed in the face because of your /d/ post history.https://desuarchive.org/d/search/filename/anistudio%2A/end/2026-08-06/
>>109476585Meds? I am not "Ani". Seems like you have some sort of personal beef with this guy.
>>109476591Keep trying. You haven't picked up a single user in 2 years but with a few more years of samefagging I'm sure at least one person will use AniStudio. You aren't being stealthy, by the way.
reference to video model is SO good man. all I needed was a 10s reference clip and a static image.https://files.catbox.moe/qh1arn.mp4
>>109476581>Unedited?yeah its just the prompt https://pastebin.com/GRVkKB6y
>>109476453why has so much artifact this is shit are you using the turbo lora experimental?
>>109476598>I eard U leik the big p-ajeet meme?
>>109476553>>109476603You're killing me with this shit, I love it.
>>109476554anon... I...
>prompt a 130cm tall girl>still comes out adult-height
>>109476553kek how did you get the barbie to stay in static doll form, thats good
>>109476643relative to something else, like as tall as the doorframe
you can add this to get better preview
I've been out since the release, so anons, what optimizations are considered lossless for h3?I know sage attention has kind of been, but what about sol attention?is the 22B version of the model kijai really the same quality as the original?any other optimization?
>>109476656Just buy an RTX Pro 6000
kino!https://files.catbox.moe/58zirn.mp4
>>109476645https://pastebin.com/dgDVmrJq
>>109476655thanks anon, how much does it reflect the actual end output?
Finally minimized pagefile and ssd thrashing, we cooking now boys.
BREAKINGusing "..." in your dialogue makes it sound all moany and asmr-like
>>109476690elipses, dash works the same
OY SHUT IT DOWN!https://files.catbox.moe/5vzfa1.mp4
>>109476685
>>109475799same/wsg/ or fuck off
>>109476705That's damn nice.
>>109476655at the cost of extra 2gb of vram
>>109476705OK that's pretty good, is it: https://github.com/simsim9-stack/ComfyUI-MiniMaxH3-PreviewOverride?
bruh, i can't get the enemy tank to shoot. it's always my own tank that shoots no matter what
>>109476099is this ai
>>109476651Thanks, I prompted "short as a child" and it worked
>>109476717try "tank in the distance shoots" idk
>>109476574>Is it true that this guy is one of H3 devs?yes "anon", ani is chinese and works in china
>>109476726>le black loud manhttps://files.catbox.moe/yb9wns.mp4
>>109476716yea, but KJ also made his version u can try that also
>>109476720i2v
>>109476732>entertainers must be as intelligent as EinsteinI don't really agree with that, intelligent people are meant to do smart shit like maths and stuff, and you want them to be clowns?
>>109476735bs, I don't believe you
>>109476574>I've heard that he flew to Japan to strike a deal or something.and he completly failed, sounds like that showing a 35 stars github project as your best achievement is a bit light as an argument to get a million dollar grant, color me shocked lol
>>109476742that guy in particular is a subhuman, all he can do is act like a monkey complain about a guy like Messi who isn't a worthless dindu.
>>109476757that said, there are many harmless "content creators" who dont act like apes in public. I would euthanize every shitty streamer that is a nuisance in public.
>>109476757>that guy is a subhuman unlike my heckin wholesome nigball kicker
>>109476757>complain about a guy like Messi who isn't a worthless dindubro have you seen how Messi acted in the finals? he's also a fucking subhuman
>>109476765shouldnt have been racist, chud
>>109476734oh yeah I'd rather minimize the number of extra repos
>>109476767I know you're joking, but there's people who really believe a white guy has been racist towards another white guy as if that's possible kek
>>109476774no, its very easy to not be racist, you should try it some time.
>>109476774>white guy
updated turbo lora is out. audio is much better. https://litter.catbox.moe/9iu2isopewmh3ptg.mp4
>>109476792https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
>>109476787>Messi is browntell me more frog
>>109476792Thanks, anon.
>>109476796>a custom node just to load a loraI really don't understand why comfy doesn't want to support the diffusers loras in the first place
>>109476799>we are white senoooooor
>>109476813wait for someone to convert it
>>109476814Messi is genuinely white though, he lives in Argentina but he's ethnically italien, your race doesn't change because you go live to a non-white country, or else Elon Musk would be a black person juste because he was born in Africa lol
>>109476825Musk is half chink half jew
>>109476820i think someone posted a conversion script, i'll try it
>>109476829Xi Long Muskstein
>>109476832https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/tree/main
>>109476825>italien>white guy
>>109476829He's half chink half niggerHis ancestors were Song dynasty Chinese maritime merchants that settled in African coastsThis was from a 23andme leak
>>109476837b-but we european like you :'(
This is 4 steps using the dual clock euler node and turbohttps://files.catbox.moe/ep9ueg.mp4
*yawn*
>>109476846>>109476834https://github.com/shuaixn/ComfyUI-MiniMaxH3DualClockSampler
>>109476844kek
>>109476792Straight up doesn't work for me, it doesn't matter what amount of steps I use it just takes even more time that the full 20 steps with no lora.Do I have to remove all the cache nodes and stuff for it to work properly?
>>1094768443 doesn't happen. All the others are accurate though
>>109476853duh
>>109476834thank you
>>109476844Mexicans don't look like this
this is the best WF
>>109476863i knew a mexican that looked exactly like that
>>109476867Check under your foreskin
>>109476748Wasn't Ani one of the first devs ever who told everyone about the potential of video generation? I have never seen anyone promoting AI videos before him. Even if Ani didn't work on H3, I think he's still partly responsible for local video revolution. He was the first to believe in it.
>>109476867Get rid of easycache
>>109476867i thought these dont stack, u only need 1?
>download a couple loras off of civit just to test the lora loader>none of them work>ask GLM why>turns out they're all for some pruned version of the model and incompatible with bf16holy poorfagerinos
>I looooooooooveee dogshit audio and fried qualityJust wait till lightx or someone actually competent makes an real lora.
>>109476883yea they do stack and kijai said to use them together
>>109476885You should train your own loras with your richfag hardware then
>>109476834this is worse than the original turbo lora
>>109476886>Just waitNOOO I CANT WAIT 10 MN FOR ONE VIDEO ANYMORE THATS A TORTURE
>>109476895lol its 10x better>>109476792>>109476846
>>109476892kek
>>109476901and this with only 4 steps and the lora trained to 500 steps so far
>>109476892>literally 1 (ONE) GPU>richfag hardwareI don't know man
>>109476878>the first devs ever who told everyone about the potential of video generation?Wow, that's some next level delusion of grandeur. All you did was make janky, dogshit gifs that no one in the industry even knows exist. I can't believe you're typing this with a straight face and not realizing the humiliation you inflict yourself.
>>109476886lol, lightx has horrible distill loras that kill motion, they can't even get it right for their own LTX model
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUIits not fully done, but to try
imagine needing poorfagtimizations
>>109476901it's very blocky for me using the same settings as the first turbo lora
>>109476912So I download the minimax_h3_turbo_4step_ckpt500_pruned_comfyui.safetensors one right?
this is new turbo lora at 8 stepshttps://files.catbox.moe/9bnw18.mp4
5mins at 20steps on a 3060 12gb for a 5sec video with audiothis is pretty nicehttps://files.catbox.moe/d4u7hf.webm
>>109476911Dunno man my guess is that at least is going to be significantly better than the stuff from literal whos we getting right now.They are on it anyways so we'll see https://github.com/ModelTC/LightX2V/commits/main/
>>109476918I think so, thats what im trying, havent genned stuff yet though
>>109476931based
>>109476931I hope you have 8x 5090s
>>109476926do u switch off the easy cache / h3 cache when u use turbo lora?
>>109476941they need that to make those turbo loras, jesus I didn't expect the requirements to be this expensive
>>109476943dont use easycache with turbo lora unless you are doing some high steps still. 20 steps with distill lora + it would prob be like doing 50 steps though if you want the quality
so normally Comfy will offload to cpu/ram, would adding a second GPU to your rig speed it up then? would it offload to the second GPU first, and only then to CPU?
>>109476941You don't?
>>109476949they dont say they are making a turbo lora, just that they are doing sparse attention along with better 8x card communication for better speeds when running it on 8x 5090s
>>109476953no
>>109476959why do you need 8 5090 cards to run minimax?? do you really need 32*8 = 256gb of vram to load everything? lol
6 steps with new turbo lora works well enoughhttps://files.catbox.moe/zmxtkk.mp4
>>109476834>https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/tree/maindon't go for 4 steps lolhttps://files.catbox.moe/vimzug.mp4
Why is it that LTX 2.3 works better the more you talk in the prompt like a complete jeet?
>>109476969its parallelism, you can have them all running the model at once in shards making it gen much faster
>>109476983so why >>109476964why can't 2 gpus do parallel?
>>109476996layers, or something
>>109476996you can with NVlink. Without it they can't communicate with each other faster than they run alone
>>109476979I smoked weed on an empty stomach once and the same exact thing happened
>>109476979 (You)8 steps still gives you terrible audio, they haven't finished cookinghttps://files.catbox.moe/zskodl.mp4
>>109476981Mechanic Jeet dataset
4060ti, ref2va with one imagewith pic related row of optimization175 seconds https://files.catbox.moe/jy3sxu.mp4with no optimizations what so ever485 seconds https://files.catbox.moe/7p6lbt.mp4
>>109477023are you using the dual clock sampler?
>>109477003how come this works then >>109476969 5090 doesn't nvlink
>>109477050cause they are developing a way to make it work. That is the whole point. Can't you read?
>>109476889proofs?
so with that lora you can gen with 10 steps (or half), so far. not fully trained but it works: didnt specify dialogue so he's mumblinghttps://files.catbox.moe/oj7vzb.mp4
>>109476899Sweet nice loop, what was the prompt?
>>109477063don't make excuses for it, gibberish while amusing is the bane of this model
>>109477063The loras makes my gens slower even at 4 steps, I'm gonna assume they are not compatible with the reference model.
>>109477033tf is that, why can't it work like normal?
>>109477073im gonna wait till it's properly finished cause with sageattn/spectrum the gens are high quality and not too long at 0.3mp.
do you need something special to use the 500 step turbo lora? it looks bad for me
>>109477078cause the audio gets negatively effected otherwise
30 steps with res_multistep and simple schedulerAnyone figured out a better sampler/scheduler combo?https://files.catbox.moe/3caocl.webm
Almost 1 week since H3 release and still no cock/pussy LoRA?
>>109477100sulphur plans to train soon. He already raised 4k of 10K goal in one day
>>109477100there's some early attempts on civit but they all likely suck
>>109476689how
>>109477095>simpleSome people say beta works better for them but it produces horrible results for me
>>109477055sorry, I did miss the point.so for now a 2 GPU build is only good for 2 separate runs at the same time then? not sure that's worth as much as a speedup of 1 run would have been
h3 50 steps does dramatically improve the motion. It feels soulful and emotionally charged.
>>109477127running two separate generations is faster. there will be more latency to have to manage the memory of multiple devices for a single one
>>109477134try the distill lora at 20 steps. Its close to 50 steps
>>109477134this but 500 steps
stardew valley, middle east editionhttps://files.catbox.moe/0n9a1a.mp4
>>109477114I set my pagefile manually to 24gb and disabled sysmem fallback in nvidia control panel.Also have this in my .bat for comfyset COMFYUI_CACHE_RAM=Trueset CUDA_MANAGED_FORCE_DEVICE_ALLOC=1set PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:64.\python_embeded\python.exe -s ComfyUI\main.py --fast fp16_accumulation --fp8_e4m3fn-text-encNot sure if that fp8 text encoder flag is helping though.euler_a a shit
>>109477142That must be something. The leap from 10 to 50 is already consequent.Also the more local models evolve, the more I hate myself for not getting a Blackwell at $10k
>>109477124Are you using the reference model?
this is why they opensourced
>>109477134does more steps improve the grainy artifacts of fast moving stuff?
>>109477163it falls off over time but yes, it will keep improving
>>109477163Image quality looks bound to the output resolution a lot.
we can fix hollywood now.https://files.catbox.moe/g9zznk.mp4
>>>/wsg/6209063Still getting the hang of the reference model. Gemma31bchan has been a big help with prompts
>>109477175t2v result: pretty good desuhttps://files.catbox.moe/ahb7p5.mp4
>>109477030sadly this setup seems to produce errors
>>109477201free vacation
Are there any options for GPU renting that could handle H3? I'm stuck on a gpuless laptop for a while and H3 is probably the first video model that seemed legit interesting to me.
mental note[96768, 2688] - bf16[96768, 8] - pruned
>>109477217Note: not some apishit service but something that would let me run a ComfyUI instance. I know people used paperspace way back in the day.
>>109477201try disabling the nodes one by one. see which one is the weak link.
>>109477217The problem with being a rentcuck is the thing you want to rent is always taken first.
A 10 second clip might have been asking a bit much...
lmao, I t2v'd a gen with bethesda in it and it got a game aesthetic, I was pretty vague, still good though.https://files.catbox.moe/mwrrxg.mp4
>>109477275gonna make breakfast and do a longer slightly more hq gen. this is fun.
>Minimax's stock is up 66% since the launch of H3they deserve it
>>109477292lmao that's awesome to see really
>>109477292but did it pay for the training costs?
>>109477292They took a gamble going open weights and it paid off
>>109477159>this is why they opensourcedi don't know why more companies don't go open source. it's basically free money, and you avoid dealing with payment processors or trying to make the model profitable
>>109477304>i don't know why more companies don't go open source.big companies like disney can sue you and say that open sourcing the model helps into facilitating the copyright infringement or some satananistic big corporation mumbo jumbo shit
>>109477304>and you avoid dealing with payment processorsthat doesn't make sense. they are already dealing with payment processors as a business
>>109477228 >>109477217maybe still vast.ai or something. ask in the non-local threads or look up that and alternatives, I'm not up to date on API shit.alternatively, give comfyui and minimax some business via the SaaS API offer as reward for giving you the option to freely switch whenever you finally want/can build your own machine? they're at least not holding your ass hostage.
any fix for sage attention doing nothing? I installed the right versions, ran the tests, no errors during gens, but the times stay the same. 3090 rtx
What turbo lora to use for the int8 convrot versions of the i2v and ref2v h3 models? I'm seeing pruned everywhere.
>>109477320>ask in the non-local threads or look up that and alternatives"Non-local" threads are mostly using proprietary models or API services. I'm interested in spinning up an instance with fully-fledged ComfyUI/whatever. I know it's not local, but it's not really "apishit" either. I'll look into vast.ai, thanks, I think I've seen it mentioned before.>give comfyui and minimax some business via the SaaS API offer as reward for giving you the option to freely switch whenever you finally want/can build your own machine? they're at least not holding your ass hostage.I'll have a look but I'm not really interested in a restricted service. If it's any good I'll try it though
>>109477292they do deserve it good for themso far, testing this model, i found that error prompting rate is low if one follows the format and quality is consistentnot sneeddance tier but at least 80% of quality indeed
>>109477195It's crazy to me that shit just werks. No hacky loras or weird checkpoints. You just give it an image and it uses it seamlessly
>>109477348Not natively enabled for H3 by comfyUI. Needs to have "--use-sage-attention" batch start flag and KJ nodes.
>>109477376Pretty sure the patch sage attention node does nothing if you use the --use-sage-attention flag. It's specifically for people who don't use that launch flag but still have sage attention installed
>>109477353if the file size of your h3 model is 22gb, then you are using the pruned one
LLMs have made setting things up so complicated that you need another LLM just to set it up for you. not all of us are advanced computer scientists
HAHAHApraise China for this model.https://files.catbox.moe/ul7310.mp4
>>109477384For your sake, I hope you're right. Cause it works on my machine.
Holy shit dpmpp_2m_sde sounds so clean and looks so good by Kijai's example.Glad samplers are now working properly so now we can gen even better looking shit.
>>109477403That's the other node, you always want that one enabled. I'm talking about the "Patch Sage Attention KJ" node which does nothing with the --use-sage-attention flag
anyone tried training with ostris experimental adapter yet? i'm gonna give it a go. but he said he'll train another better version
>>109477402>oh no my pronouns were ignored
>>109477357Maybe this? https://www.thundercompute.com/cloud
>>109477369I notice that if there's an object at say 9s that is entirely new to the scene at that time but is say hidden, like glasses for example, even though they are not being acted upon (unfolded and placed on the head) they appear in the scene before the 9s mark in the persons hands.I'd not noticed this before. Nifty.
Hope you fuckers are not genning CP with the model because it's very capable
>>109477402version 2:https://files.catbox.moe/gk35st.mp4
>>109477393if i could do it, you should be able to as well
>>109477357>I'll have a look but I'm not really interested in a restricted service. If it's any good I'll try it thoughSure, I leave it to you to try. If you ultimately toss them a few bucks, IDK if it will incentivize future models but they certainly deserve it for this.
>>109477391The convrot int 8 is 33gb, so no. There really isn't a turbo lora for that yet, why?
tried timestamps. almost fully worked, I could refine the prompt a bit.https://files.catbox.moe/65phjn.mp4
The new lora is fucked at whatever settings people were saying to use before, it's seems ok at a lower strength? 1.2/1.3 is way too much
>>109477498i think you're looking at the text encoder, not the diffusion model. the int8 for the diffusion model is about 22gb
>>109477503Most of the time I'm retarded, but not this time I think.
>>109476485you didn't even do it right, retard. You're calling the character <Picture 1> instead of <Subject 1>
>>109477501Do you want to know how I know you're a jeet?
>>109477513how is he retard if it worked?
>>109477509oops, sorry. download the pruned versions, there is no reason to use the original ones unless you are training things
>>109477421The prices look nice at first glance, thanks.>>109477466It's kinda fun seeing Comfy grow so big that his corp is running their own SaaS and funding the big models. I mostly remember him from the SD 1.5 days when he felt like an asshole shilling his weird UI but the dude really was serious about this stuff. And yeah, if the comfy service is not anal about horny prompts I'll probably go for it.
>>109476147>Maybe specify "no dialogue" in your promptTried, doesn't work.
holy shit, t2v not even reference audio. I want a data list of known characters just out of curiosity.https://files.catbox.moe/mbvnro.mp4
>>109477541>t2vyawn
Anon you did get claude to make you a minimax prompting tool that can connect to your LLM that's running on your network so you can greatly increase the quality of your H3 gens, RIGHT?
>>109477543r2v is my favorite model by far cause you can plug in any character or any audio, im just testing both models
>>109477547my pc is the most powerful on this network, I do have a spare 12gb card but any LLM that can fit on it won't be worth a damn
woah 36 stars nowholy shit
>>109477550Might be able to use a quant of gemma 4.
With video continuation, is the meta to cut and re-encode the video with a lower frame rate, right?I tried continuing a 9-second gen last night, using the video + audio track as reference and it took 1400+ seconds. The result was good but damn that's a long time.Then I cut the video length in half (trimmed the start) and encoded at 20 fps (could've probably gone less), tried running the same prompt with the re-encode as input, and it took ~700 seconds.
>>109477543the video is in the style of a cinematic Star Wars movie with Yoda as the star.0 to 5s: Yoda is standing beside a tree in front of a pond, in the forest. Yoda says "blacks problem yes, remove blacks you must do!"5 to 10s: close up of yoda who says "otherwise, bike steal, they do." the camera zooms out and a regular bicycle near him is force pushed away by Yoda.https://files.catbox.moe/la18am.mp4
>>109477534exactly why Ani is destined to win. comfy still won't fix shit that people were complaining about for years.
Nvidia_RTX_Nodesgonna try this with 0.3 or 0.4 gens
How do I prompt petite body but with huge breasts? Every time I do this the model gives me a landwhale
>>109477569Who win what?
>>109477601UI wars
Can one of you cucks explain to me how the cloud providers will be able to sustain their video generation?Sora already had to shut down because it wasn't profitable. Now that Minimax H3 is out, and it's outstanding, why would anyone pay for shit like VEO 3 and Kling? And wan, what are alibaba going to do there? Absolutely nobody will pay for wan at this point.
>>109477609sora 2 gave away 30 generations free a day to millions for months
>>109477609>>109477614also sora 2 was making a bit on having companies pay to license with them. But they got denied. And what's funny is now Disney has teamed up with bytedance to use seedance
>>109477614Sora 2 wasn't goodIt was worse than H3/Seedance 2.0
>>109477609Whoever's able to monetize NSFW saas vidgen despite is gonna make billions.
>>109477620ehh... it still had magic that even seedance 2.5 does not. it knew SO MUCH. It was clearly a BIG model
>>109477547Close, I made claude make me a workflow for an interactive avatar connected to an LLM. I send it questions or requests and it sends me video answers, it speaks and grabs props. It uses a picture for reference (and it can look at its own clothes) and an audio reference for voice.
>>109477620H3 is amazing but come on, Sora was the best video gen model BY FAR quality-wise.
>>109477624you technically can't because of the agreement you need to have with the company but this is reality so there is probably some Russian telegram you can gen porn on
if on nvidia, try this as you can gen at lower res (faster) and get better outputs, essentially dlss for comfythe nvidia file was only 400mb
>>109477609>Sora already had to shut down because it wasn't profitableit wasn't profitable when compared to using the same hardware for making and selling tokens for programmers
>>109477639*change scale to 1 for 1:1 video
Goddamn, it's really good.https://files.catbox.moe/bxj2nv.mp4
>>109477624you cant monetize nsfw vidgen because of payment processors. crypto was supposed to solve this but it just became a meme instead
>>109477556Continuation doesn't work properly. Even when you very clearly state in the prompt that the target video is a direct continuation of <Video 1> and set all the labels and parameters right, it still won't be a seamless continuation and will feel like a jumpcut. Think you just have to manually save the final frame and use that as the first frame, kind of like with wan chained continuation.
>>109477645Proven false because H3 is a very strong model compared to frontier video models and can run on shitboxes, while the best model /lmg/ can run still trails frontier LLMs
>>109477514You've made me curious.
>>109477547no i've been writing this thing for almost 2 days now
>>109477661>crypto was supposed to solve this but it just became a meme insteadwouldn't it "kind of" still work here? not that I'm going to be doing this, but in principle it should be possible to get some very small crowd of cryptofags to send you their coins no?
rtx upscale at 1.0 test, same yoda:https://files.catbox.moe/16kiy7.mp4
>>109477547I just use SillyTavern and dump the whole two prompt instruction docs into the context and then tell it what I want after that.
>>109477686>>109477686
>>109477609What percentage of chinks has a good GPU? Last time I heard from a chink friend, most of them still goes to gaming cafes. Also third world browns.
>>109477678>very small crowd of cryptofags to send you their coins noThat wouldn't be enough to justify the costs probably. Most people are too scared of crypto.
>>109477678DLsite does it so it could work here, too.
>>109477627Stop sucking Emliy Youcis's cock. Even she is on board with H3 being the best
messed up my wf or comfyuiit is taking double the time to gen a video with minimax h3!tried installing sage 2.2, reverted back to sage 1 but comfy is performing worse now!need maybe a better workflow for ref! anyone care to halp please?
>>109477534definitely put in a lot of work since.
>>109477748for references, video input can make it a lot longer, pictures and audio not so much
>>109477664yeah it runs but on medium-high local hardware you can gen a 10 second H3 video or about 50k tokens of current best-easily-runnable LLM, the latter feels a lot easier to package up and sell for a given price plus you get free training data off the input people send you, which potentially is novel fully human-authored text but at the very least is their AI-generated but curated codebase which they as a human assessed as being worth continuing to improve and work on. i generally assume that's why the OpenAI/Anthropic personal coding plans are so well priced relative to the API $/token businesses have to pay rather than them just pricing solo individuals out of the market; we're worth it for the authored input variety
>>109477666Bollywood slow mo.
>>109477159>>109477304>i don't know why more companies don't go open source.they only got a stock jump because they published literal fucking global sota video gen model (given seedance got censored/nerfed), not really because they went open source, although that did help since it guarantess that it wont get censored, that people can fine tune it, that it gives you full privacy etc
>>109477952I gotcha, makes sense I suppose.
>>109478108truth is there was a pump yesterday, check the indexes
>>109477614>sora 2 gave away 30 generations free a dayamazing
Hi anons. Anyone tried running this on 3090 and 32gb ram? is it painfully slow? is 5090 and 64gb ram the minimum for good time with minimax H3? I want to atleast be able to watch youtube ishowspeed while i wait for the prompt to finish. Maybe also play a match of Dota 2 while i queue up some prompts.