Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109471487https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
https://files.catbox.moe/u6ou7n.mp4we must be better men and not stack cache nodes!
What's with all the troll bakes recently full of off-topic links?
Reposting the question here, it's an interesting one>>109472926>Why is no one talking about how to improve the quality, only about how to gen faster?
>>109472946btw I haven't found anything better than res_multistep & simple, I asked Codex and it said that matches up fairly closely with what minimax recommends.
>>109472944why would the original OP be asleep at this time? does he live in india or something?
Blessed thread of frenship
I will NOT re-install cumfartui completely all over again just to run the latest model potentially faster.. No.. Please don't make me reinstall dependencies and redownload fucking torch again.
>>109472972Euler
>>109472936thanks for the bake
>mfw Resource news08/05/2026>Inline Studio v1.2.62 - Minimax H3 Lora training still onlyhttps://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3https://huggingface.co/Kijai/MiniMax-H3-TAE>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inferencehttps://github.com/6somehow/DAC-SPADE>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generationhttps://github.com/yizzz927/CAPE-T2V>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusionhttps://github.com/jd-opensource/JoyAI-Video-Edit>ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMshttps://github.com/YangYangGirl/ParVL>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diethttps://huggingface.co/JamesZar/OliveGemma-3B08/04/2026>stable-diffusion.cpp adds support for MiniMax-H3https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES timehttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editinghttps://github.com/IntMeGroup/MIEScore>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videoshttps://rathgrith.github.io/PeCA>Kandinsky WM 1.0: A family of models for Physical AIhttps://github.com/kandinskylab/kandinsky-wm08/03/2026>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>109472986>re-install cumfartui completely all over again just to run the latest model potentially faster.Source ? Is it really makes it faster ??
>>109472965you make it as fast as you can so you can iterate faster, the faster you iterate the better your workflow gets, the better your workflow gets the better your gens get, then you focus on quality.
>mfw Research news08/05/2026>SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrievalhttps://arxiv.org/abs/2608.03120>HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Planehttps://arxiv.org/abs/2608.03422>DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformershttps://arxiv.org/abs/2608.03082>Can T2I Models Draw from the Right Frame of Reference?https://arxiv.org/abs/2608.03357>Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editinghttps://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0>Self-Supervised Representation-Guided Generative Dataset Distillationhttps://arxiv.org/abs/2608.03218>Latent Reward Registers for Diffusion Preference Alignmenthttps://arxiv.org/abs/2608.03929>UniWorld-Design: From Pixel Generation to Layer-Native Designhttps://arxiv.org/abs/2608.03971>MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Bindinghttps://arxiv.org/abs/2608.03708>RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editinghttps://arxiv.org/abs/2608.03059>Efficient Video Dataset Distillation via Cluster-Guided Prototype Blendinghttps://arxiv.org/abs/2608.03269>Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfoldshttps://arxiv.org/abs/2608.03135>TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Modelshttps://arxiv.org/abs/2608.03057>Adaptive Two-Stage Visual Token Pruning for Efficient Inference in VLMshttps://arxiv.org/abs/2608.03112>Enhancing VLM Reward Models Through Structure-Aware Fine-Tuninghttps://arxiv.org/abs/2608.03875>Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understandinghttps://qwen-3d.github.io>When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardwarehttps://arxiv.org/abs/2608.03649
>>109472953loool, so which cache node is the best then? still spectrum?
>>109472986just pull it bro
>>109472965because if you can gen faster without altering the quality too much you can go for higher res with acceptable rendering times
>>109473011that anon's custom settings he posted for the cache node might put it slightly above spectrum
>>109473022the time is the same or is it faster?
>>109472986>reinstalljust update it dumbass
sometimes Minimax makes the character talk gibberish just before it says its supposed sentence, how do you prevent that?
redgjsieorjtg aijaw >>109473049 not sure what you mean.
>>109473054kek
>>109473027spectrum is definitely slower.
>>109473007thisdiffusion generation is a spaghetti at the wall situation
Any news about a 4 steps turbo lora? It's been one day already!!! >>>/wsg/6208816
>>109472989that does seem better on a quick test, but I think I'd need to compare like 10 vs 10 generations to be sure.
>>109473064no way fag, spectrum was miles faster for me. but i THINK it's lower quality. i need to run one more gen (so it'll take like 5 minutes longer) to be sure though.
>>109473075>a 12.5gb lorais he serious??
>>109472972>I haven't found anything better than res_multistep & simpleSame, I tried a lot of different settings but res_multi/simple always works.
>>109473072sacrilege!!
the cache node might fuck with fine detail and prompt adherence
>>109473072Does it work the other way around, too?
Can you handle my workflow anon?https://files.catbox.moe/eb2xd3.mp4
>>109473123>mightfigure it out
>>109472180Oh, you can make silent videos, but it's still gonna create an audio track in the output file, so you can't really turn it all the way off.
https://files.catbox.moe/b4b6ab.webm>>>/wsg/6208834
>>109473128?potionseller refusing to sell his strongest potion
tried to train innie lora for h3 results were kinda meh so didnt upload to civit but whatever here u go anons. dont @ me if it underperforms.https://gofile.io/d/9fvLNJ
what the hell is this nigga
time to test out this newfangled "spectrum" you guys been talking about
>>109473049I think it has to do with the video being longer than however long it takes for the to say the specified dialogue. You either need to keep them talking long enough to match the vid length or clearly specify in the description what they are doing before or after speaking.
Does anyone here have a workflow for using a reference image (a person) and copying the subject onto an existing video?
>>109473170someone is using your comfyui from outside your network. you know basic network security, right?
>>109473175test it out? You're already on it!dohohoho!!
Is there any upper limit for how long the reference audio for a voice should be?
does /ldg/ think minmax is good with 3dcg?
>>109473182I was lazy and tried to just throw a 2 minute collection of a character's voice and it was really fucked up. Was a lot better after I took just one sentence.
>>109473170maybe this can help you, I don't have this anymore since I went for the --disable-api-nodes flag thoughhttps://github.com/Comfy-Org/ComfyUI/issues/12618#issuecomment-3957464383
>>109473189sure?
>>109473192Oh damn. So just throwing more material at it to improve the quality isn't viable then.
>>109473180Why would a default install of comfy allow this to happen? No, I don't know network security, I am basically a consumer.
>>109473181`_`
>>109473194Thank you, I will try it out later.
>>109473203if you close your web browser / comfy tab while the web socket is open you're going to get that error
>>109473203uh oh, good luck
lmao easycache raped the prompt adherence so hard, shit looked like a robotic powerpoint presentation
>>109473198tits not JIGGLY and BOUNCY enough
>>109473218Why would you use easycache over spectrum?
The reward feels better because I waited so long to get a good local video model, hooray!!https://files.catbox.moe/iz8b93.mp4
>>109473227trying botheasycache is kinda decent for previs to see if the prompt at the baseline does the required stuff
>>109473189https://litter.catbox.moe/jn103e.webm
I really too indian and dont have creativity to do this.How to stop being indians bros ? I wash my butt with water at least...
>>109473229GOD HE'S LITERALLY ME FR FR>>109473159man you just reminded me this character even exists, gonna make a vid of her shaking her ass or something while telling a joke.
ok imma need the training config for that sweeny lora posted a few threads back, that shit is too good
>>109473240give an image to a LLM and ask it to give you 10 ideas, it can help
still no sora though
>8s 0,4mp ref2v>10min>0.5mp>14m30sgod damn
>>109473257>ref2vwhat all are you feeding into it?
>>109473239based, Daisy is the best waifu in Mario's land
>>109473257
>>1094732658s video
>>109473278what resolution is the 8s video? because 8s ref on 8s gen is equivalent to 16s if the res is the same
>>109473246>sweeny lorayou can use a reference of her though, works well
btw steps matter, 40 steps will take 2x long but if you're going for a best quality video you will see a big differenceneed speedup loras!
Is there an ideal image dimension for the reference images?
>>109473284yeah it's the same res
for me it's 100 steps easycache
>>109473298obviously, we're coping with 20 steps and it's pretty good already
>>109473298ive been coping with 14 steps, is 20 steps a big step up?
>>109473298>>109473312Why don't they release the turbo lora at launch?
>>109473315i haven't gone as low as 14 steps, you should just test it out. but doubling the steps you will see a big difference, i guarantee
https://github.com/Comfy-Org/ComfyUI/pull/15334>Adds suppport for int8_convrot quantized VAE for MiniMax-H3, tested that the VAE quality looks fine and it's about 1.5x faster.meh, the vae decoding isn't that slow, I'll pass
>>109473300one million pixels!
>>109473315>ive been coping with 14 steps, is 20 steps a big step up?go for 20 steps + spectrum, you'll get something as fast as 14 steps but a better qualityhttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
>>109473334i've already been using spectrum with 14 steps lol
>>109473326every second counts
>>109473339based.
>>109473326shit go hard when this node light up u know u about to get lit
>>109473298>btw steps matterJesus someone rape these newfags already
>>109473356anon... it's the newfags i'm trying to educate. if you haven't noticed we have an influx of like a billion of them every new model release
>>109473334okay ill try out the spectrum meme. what are the best settings?
https://files.catbox.moe/e14qfm.webm
>>109473375the ones you get when you load the node
>1 hour wasted on failed attemptskino
>>109473387thats why we need the turbo lora so we can iterate faster
https://files.catbox.moe/9jq2rj.mp4
isn't spectrum capped at 20 steps?
>>109473409no you can go for any steps that you want, but going for too low won't really work, spectrum will consider that all the steps are important
https://files.catbox.moe/syp9o3.webm
>>109473386lol wish /adt/ had more of this type shit
https://files.catbox.moe/0q609f.mp4
>>109473370Can I still rape you tho
>>109473448>>109473424he just likes the smell of an unwashed bellybutton
>>109473454NTA but you have my permission.
>>109473470can't be rape if there's a permission though
https://files.catbox.moe/wuis74.mp4
>>109472999>>109473010thanks!
>>109473179it's a shot in the dark right now. And I think it depends more on prompting than workflow.
>>109473479I take it back then. You do NOT have permission.
>>109473480Now this is what a peak gen looks like. Everyone else can just delete H3 now.
H3 actually doesn't output exact duration. Sometimes you get trailing 0.35 seconds or something and it can fuck your alignment up if you chain multiple together. FYI
>>109473507the reason>max(5, round(a * b)) + (5 - (max(5, round(a * b)) % 17)) % 17
would be a great advertisment video for Minimax lolhttps://files.catbox.moe/18inkw.mp4
>>109473507it's 17k+5, the math is in the note connected to duration
are abliterated text encoders just snakeoil slop?
I don't know how to run MiMax on 12gb vram and at this point I'm too afraid to ask
>>109473524yes, it's not censored to begin with, the heretic guy just blanket applies uncensored slopping to every LLM out there. it just makes it stupider.
It can't deal with reflections properly, or I cannot make the prompt clear enough make it work.
Having a hard time trying to prompt barefoot footsteps sound. H3 always wants to make it sound like high heels
>>109473529you download the int8 model and you let comfyui do the automatic RAM offloading for you, if you have OOM, try those flags>--vram-headroom 1>--disable-pinned-memory
>>109473494https://files.catbox.moe/h77iw4.mp4
>>109473532it definitely feels like some things were omitted from the text prompts in the dataset, but no videos outright culled from it either, leading to some concepts just... happening when the model feels like it
>>109473535plod, thump, uh.. thud
>>109473534Prompt issue
>>109473170what are you running it on, are you doing it remotely? looks to me like you just got owned. You should be connecting to comfyui remotely as if it was a local IP through SSH or something and not just opening a port. Now the question is how fucked are you?
>>109473535try more descriptive, like "to the soft padding of feet on carpet"
>>109472883alright, quick update Copied these settings from [spoiler]reddit[/spoiler] and now I'm getting much faster generation time. Did some comparisons and the quality hit was neglibile, almost like a variation of the same thing. No H3 Cache, 0.3MP @ 16:9, 10 seconds: 647sH3 Cache Settings from last thread: ~450sH3 Cache settings fro Reddit: 372s
>>109473203are you attempting to run comfyui through 2 different web browser tabs?
>>109473534the gunt is fucking disgusting
>>109473553if you pulled master these are now the default settings for the node.
it can do ASMR toohttps://files.catbox.moe/st0f5h.mp4
>>109473542What's more likely is that the LLM-hallucinated prompts aren't that accurate
>>109473561can it do duke nukem asmr
>>109473561what kind of prompt for that? i like it
no audio so it doesnt fry your brain
>>1094735740:00–0:02Mickey Mouse sits at a tiny ASMR desk in a cozy, softly lit studio, positioned extremely close to a binaural microphone. He gently taps the microphone with his white-gloved fingertips and whispers in his unmistakable high-pitched Mickey Mouse voice:“Gosh… you made this?”0:02–0:05He slowly leans closer, maintaining cheerful eye contact while softly crinkling a sheet of legal paper beside the microphone:“Well, I’m gonna have to sue ya…”0:05–0:08Mickey pauses, gives the camera a playfully threatening smile, then whispers:“Ha-ha! See you in court!”He finishes with his distinctive “Ha-ha!” Mickey laugh, followed by two delicate microphone taps.Exact classic Mickey Mouse appearance, voice, mannerisms, white gloves, red shorts and yellow shoes. Intimate stereo ASMR audio, soft whispering, gentle paper sounds, detailed glove tapping, no background music.
>>109473579catbox? curious what the prompt was
>>109473561AAAAAH DISNEY SAVE ME
https://files.catbox.moe/65o65y.mp4
>using params from plebbit >thinking spoilers work on /g/ newfag holocaust when
>>109473583Then ask for prompt and not catbox.
>>109473590i do not like the way that miku looks
>>109473581thank you
>>109473594walls of text are unsightly
>>109473583
>>109473604Damn H3 threw in the jiggly boobies for FREE?
i... last one, i promisehttps://files.catbox.moe/nsxltn.mp4
>>109473608
>>109473559ah... fml
>>109473610kek
>>109473579>no audioI want the audio.
>>109473610hi Lodestone!>>109473601I had a better migu on the first try but she appeared before he asked for it, I don't want to wait 10 more mn, take it or leave it :(
>>109473591>NOOO YOUR SUPPOSED TO SPEND ALL YOUR TIME ON ONE BOARDa bloo bloo bloo gramps, a bloo bloo bloo
used to be that redditors would steal content from 4chan. now we steal content from reddit. Grim.
what settings should I tweak on spectrum to match the speedup done by easycache's default settings? i'm kinda clueless what any of them do
>skip steps: 100%
>>109473610cute dog
>>109473632and discord, are anons unable to make kinos anymore?
I have no idea what im doing with my prompts. Granted im using text-only so this is probably not as trivial, but honestly even with the guide for prompting they supply I feel fucking retarded.https://files.catbox.moe/ia1b4n.mp4Is there any particular short tip that you have?
>>109473635>looks at blank result>"ah, it's exactly what I envisioned in my head"
>>109473632>redditeverything is reposts from trooncord
>>109473507FYI for multiple shot chads out there.
trying to get h3 prompting down, there are two guides in the docs. should I be using one over the other? I imagine one is for exact super specific prompts and the other for more general vague prompts ?
>>109473645one is for the FL2VA modelone is for the REF2VA model
>>109473632>>109473643almost as if ledditors and trooncord users are also using 4chan and are crossposting their videos or something...
>On Monday, Tokyo police announced the arrest of a 32-year-old man from Himeji, Hyogo for using generative AI to create nude images of female track and field athletes — some of them as young as 12 — and posting them to social media groups. “It made me happy that I could make people happy,” the suspect reportedly told police.>Investigators say he created around 300 images and posted roughly a third of them, some at the request of a 17-year-old member of the same group, who has also been referred to prosecutors. One victim is now in her 40s; the source photo used to create her deepfake was taken when she was in her first year of middle school.>Japan has seen a string of deepfake-related arrests over the past year, including a man accused of selling more than 500,000 AI-generated sexual images of celebrities and another accused of creating deepfakes of more than 260 female celebrities. Each case has relied on a different legal theory, exposing the same underlying reality: while generative AI has transformed how these images are made, Japan’s legal response is still built from laws written for something else.https://www.tokyoweekender.com/japan-life/news-and-opinion/man-arrested-over-ai-generated-nude-images/
>>109473648ahh that makes sense, ty anon I was very confused
>>109473651>500,000 AI-generatedjesus this dude never sleeps or what?
>>109473651>and posted roughly a third of themwell, duh
>>109473651just don't post your shit online, it's that simple
>>109473651>sharing the gens online>using real people
>>109473651whats the moral of the story?
>>109473651>Has go into games machine>Proceeds to share his waifu with others like a cuck in those hentai>get arrested.Deserves the rope.
>>109473651that's quite a collection
anyone been playing around with steps in minimax? 20's the default but how low can you go with negilible losses?
Do seeds affect gen time?Same settings prompt but a +1 difference in seed
>>109473651>It made me happy that I could make people happyHe didn't deserve it :(
>>109473651>>>109473653 >>109473655 >>109473657 >>109473658 >>109473660 >>109473665 >>109473670least obvious botted replies
>>109473676beep
>>109473676What would be the purpose of the bot?
Why genning with Minimax cause my GPU fan to load 90% while genning with LTX only like 50% ?? I thought it only uses VRAM ????
>>109473534>1slag
>>109473651based/ldg/'s ultimate boss
>>109473676>anon posts headline about ai crime>anons respond to it cuz the issue here was obvious>Umm.... le botting?!?!
why does non_diegetic_music pause when ever people talk, how to stop that shit because its really fucking annoying and stupid. i'm not fond of this boomer prompting since i've had better results just prompting it normally except when the subjects talk, then its better to do something like. the man says: (S1) <d> [English] text. But then again I've had some success with just literally. The man says "text"Audio: sound happensalso seems to work so what the hell gives? Its not at all easy if it doesn't even work as described in the documentation. I'm gonna try some different methods, perhaps i can be structured with action in the video as time stamps and then followed by new paragraph with time stamps for dialect followed be another section with timestamps for the audio.
>>109473684>he doesn't pin his fans to 100% when generatingenjoy your broken gpu
>>109473676you must have autism or something lmao
>>109473651Nah. The Tokyo police are pervs too. They just want an excuse to talk to the purported "female victims".
>>109473672I'm using 15 steps just to save some time, 10 is too low aside from a rough test of a prompt. And it's a 480p as well, only on a 5060ti 16gb so not blessed with the speed of the better cards. I mean really 20 steps should be used. But this early on with still trying out prompts etc and needing a few gens done per prompt I take the hit on quality. To me it's about the content generated not just the visual quality anyway
>>109473645>>109473648One is base that works for both, for things like camera control.
Is the anon who made the rimming videos earlier still here?
>>109473695probably try something like 'music plays uninterrupted as'. it'll halt action states at shot transitions, too.
>>109473551I am just using the .bat from the portable version. What's the worst that can happen to me now?
>>109473653Easy to automate a SD.1.5 script on Comfy or Auto1111 on a good GPU and make thousands of images per day.
https://github.com/SandAI-org/MAGI-2-previewhttps://sand.ai/blog/magi-2-preview>114B total parameters>no video showcaselmaoooooo
>>109473735look at the requirements
>>109473674.
>>109473710you're here, so yes
>>109473742No, I had a question for them
>>109473674if you use the cache nodes (not spectrum) it's possible one seed skips less steps.
Sigma chads, what settings are you using, and are you getting a noticeable improvement to audio or visual quality
>>109473743rimjobanon here. what's your question?
>>1094737598/4, yes.
>>109473735Wut? What a disgrace to use that venerable name with no result to show for it.
>>109473284So I can lower my ref video res and speed things up?
https://files.catbox.moe/w8cvp7.mp4how do I STOP it adding sound effects?I told it:overall_soundscape:No ambient, environmental, or physical action sound is present at any point; the wind, footsteps, impacts, and explosion play silently. The theme song is the only audio in the entire video.but it still added punches...
>>109473743>them
>>109473735>we've made a 100b parameter model>that activates 6b per/tk>and it needs 8 HOPPER gpus https://www.techpowerup.com/gpu-specs/nvidia-gh100.g1011>all of themwat
if i am not running blackwell, should i still use nvfp4 or should i just use q4?
>>109473266Why is every model so bad at placing the smoke for cigarettes?
>>109473772don't you know the first rule of prompting is to avoid saying shit like>Don't do X
>>109473735>>109473775Imagine where we would be if AMD was on the table for this shit
>>109473772Sound prompting seems to be pretty buck wild. I have gotten it to add sounds only inconsistently despite matching the prompting guide. It almost never misses dialogue but sound effects just don't show up half the time, or more. You could try "n/a" in the overall soundscape field but it seems to be a bit voodoo.
>>109473777sometimes smoke can travel up the cig like that
>>109473775How does a 100B model require 8 80GB VRAM GPUs?
>>109473513how long did it take you to make that? What was the process?
>>109473784>if AMD was on the tableAMD is a controlled opposition
>>109473735114B is a small model in /lmg/
>>109473763I was wondering how you inserted a video into the R2V node, but it seems like you used comfyhelpersuite after searching a few threads earlier.
>>109473579is this overwatch?
>>109473818fortnite... overwatch... same shit
>>109473558me gusta
>>109473651>posting them to social media groupsdude fell for the modern day honeypotsI avoid those places like the plague
>>109473806NTA but in default comfyui there's a load video node and another node for splitting the video into image and audio. Use both and then route them into the respective slots.
What node do you use to load a video in r2v? the comfycore load video node has a grey output but R2v takes a blue input.
Please chinkmootBanish these newfags to the shadow realm Thanks, Anon
/ldg/ got chinese cultured
>>109473842Video Helper Suite, custom node.
>>109473849Huh? Isn't that Kijai's node pack? Why would Minimax R2v require this?
>>109473759pixart sigma?
>>109473853good question
So I heard h3 makers were going after anyone making nsfw and generally banning that usage, which means places like civitai will not allow it either.I guess that means the model will be doa long term for "complete" nsfw?
>>109473853it doesn't. You can use Get Video Components. But the custom node is better.
so then, after a few days of everyone niggerrigging and niggergenning, which scales hardest against our s/it? runtime or resolution? getting hit harder trying to gen at 1mp less than 10s, or 10+s at 0.5mp? where's the middleground?
>>109473771you can lower res, lower fps too if you want
>>109473864looks like we're fine, there's already a lot of nsfw h3 loras on civitai and none of them are nuked
>>109473842Load Video -> Get Video Components
>>109473803that's captialism for ya
>>109473881Oh nice, then what happened, last I read they were on a nsfw hunt?
>>109473853gee, anon, I don't know. Maybe the colors and input names mean something.
>>109473892>>109471258
>>109473892apparently civitai got a loicense from em
>>109473892only on hugging face so far.
>>109473899>tfw you remember you have free will
>>109473884TY anon this is what I needed :D
kino incoming
>>109473909>you have free willyou don't
>>109473553>>109473559fuck. Alright, fine, I'll add it to my testing. Can I add it through the manager, or do I have to git clone?
If you want quality while saving vram, use the gguf quants of H3.The Int4 variants gave garbage outputs (like really bad, as if the conversion didn't work correctly). Int8 may also have this issue.The gguf quants are slower, but I actually get decent outputs. After applying sageattention 2.2.0 and Spectrum with conservative settings, I cut gen times down from 15~mins to 8~mins on a rtx3060 12gb and 24gb of system memory for a 5s I2VA clip at 0.4mp.Using H3 Q4 quant and nif0's qwen3vl 32b ultra herectic IQ2_XSS with his forked gguf loader that adds support for the mmproj vision file. Also using molbal's gguf loader for the main model so I had to rename nif0's gguf loader custom node folder to tell the difference.Standard UNet GGUF loader wont work here as it's using nif0's node, using the dynamic vram version using molbal's node instead.https://files.catbox.moe/uxb7x1.webm
>>109473929>5s>480p>ggufsnigga please fuck you talm 'bout
>>109473791n/a didn't work either. Maybe it's just not built for non_diegetic sounds only.https://files.catbox.moe/gk65xc.mp4
>>109473929Weren't you that fag claiming LTX still has a usecase or am I mixing up my faggotposters
>>109473929mmm no
>>109473804is not like we can use multiple gpus in ldg
Holy shit a turbo lora is already herehttps://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lorahere's a lora strength comparison, it's 8 steps + euler + simple for all
33 mins for a 20s vid fuck outta here lol
>>109473951fuck you it takes me that long for an 8s one
>>109473950based.... so based...
>>109473929can you post an example that has more movement than a flapping mouth and slight arm move
>>109473950Still a bit undercooked? The hands...
Ok, the leddit settings for the h3 cache node definitely help gen times for me. 38% faster in one trial. There might be quality loss, but it's hard to tell - as in, it's not noticeable compared to ordinary seed variation.Comparing:INT8 convrot -> minimax sage attn patch -> spectrumINT8 convrot -> minimax sage attn patch -> cache node.3mp square, 15 secondsCache: 350 secondsSpectrum: 482 seconds
>>109473950wtf? how does it look better than no lora?
>>109473975>how does it look better than no lora?because it's at 8 steps for all anon, regular minimax at 8 steps looks like shit
bros I didn't even ask her to lick her finger...she just has a mind of her own!https://files.catbox.moe/2mj4wt.mp4
>>109473772>overall_soundscape:overall_soundscape: N/A
I have another R2V question.How do you reference the audio from a video reference?The prompt guide says:>2.5 Visual and Audio Tracks from the Same Reference Video> <Video N> and <Audio N> are numbered independently. Each index indicates only the label's order within its own category and does not encode a pairing between the two categories. The same reference video may therefore correspond to <Video 1> and <Audio 2>; different indices do not prevent them from coming from the same source asset.>An ordinary reference video does not create <Audio N> merely because the file contains sound.>An <Audio N> definition primarily states the audio's role and does not have to name the <Video N> it comes from. State the shared source only when needed to remove provenance ambiguity, for example:> <Video 1> is the source video for the target video edit.> <Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.But as seen from the node (pic related) there is an input for "ref_video_audio_0" which increments upward.If I reference <Audio 1> with pic related, will that reference the "ref_video_audio_0" input, or the empty "audio_0" input? The documentation doesn't make this 100% clear.
>>109473950>>109473975i still dont understand how turbo loras work. why dont the model creators just make it turbo from the beginning, is that what zit did?
>>109473988>why dont the model creators just make it turbo from the beginning, is that what zit did?yeah and it makes the model impossible to train
>>109473988I didn't for a long time and I was very skeptical, sounded way too good to be true
>>109473898>>109473900>>109473903Thanks anons, well good thing then.Big websites will get their licenses and people will do whatever they want.
>>109473950Perhaps my 10 second gen dreams will come true after 3 days(?) Shit is moving fast
wtf, to make it loop you just have to tell it to loop.>Preserve a seamless loop and stable camera.
>>109473772>The theme song is the only audio in the entire video.use non_diegetic_music: None interrupted fast paced techno drum beat with bassline.Or what ever.
>>109473985Prompting starts from 1. The node slots is just coded to labeled from 0.
>>109473988Better to have the makers focus on the full fat thing then everyone can do whatever optimizations and quants they want after
>>109473985from my understanding they are all in the "<Audio>" queue but all the ref_videos are added to the queue first. so the ref_video_audio_0 is Audio 1 and ref_audio_0 is <Audio 2> if both connected. If no video ref_audio_0 is <Audio 1>
>>109473950https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/discussions/1#6a73cf519b0aee71a9a71bcd>Nah, this is just an experimental demo, not comyfui compatible yet. Will add comfyui support once the full training finished.motherfucker I wanted to try it...
>[WARNING] lora key not loaded: token_refiner.blocks.1.attn.qkv_proj.lora_A.weight
ok so turbo lora at 4steps euler/simplegibberish
>>109474019it's in the oven, people are working on it, relax
>>109474019>not comyfui compatible yet.bruh.
>>109473950
im going back to ltx. h3 is not mature yet
>>109474028>it's in the ovenhe's posting it on huggingface though, so he wanted us to see it and try it
>>109474038see >>109474019explicitly says it's happening but it's not readyfrom this we learn that it's happening, see
>>109474032lmao why the fuck did the baby getting thrown 3 seconds in make me laugh harder than your actual original gen
>>109474037>im going back to ltx. h3 is not mature yet...you sick fucks!
>>109474032holy sovl
>>109474032holy fucking white teeth
>>109474037He says while everyone is posting kinos left and right that not even the top cloud models could produce.
>>109474032>i'll give you a 5090 if you spin around and throw your bab-
>>109473950wooooooooooow>>109474032kekk. more lol
>>109474054you're inviting a british jokeoh hell I'll do it myself; >t. british
>>109474032that's how the game should have been desu, imagine raising someone else's kid, yikes
>>109474032>When you don't get child support for McDonalds
>>109474071kek to the extreme
>>109474019someone get claude to convert that generate.py into a comfy node.
what are you genning with H3 RIGHT NOW?
>>109474032>she spinned so fast she duplicated himkek, the laws of physics are weird when we're getting close to the speed of light I see
>accidentally wrote fish instead of fistoh man
In case flux3 isn't completely shit, can you technically use the output of minimax to teach it concepts?
>>109474084tacky mind control fetish nonsenseno other model has ever known how to do this out of the box even if it's all b-movie tier acting
>>109474084more migu >>>/wsg/6208856
>>109474019i used it just now but the quality sucks, guess that's what they meant by comfy support?
>>109474084kinos
>>109474103>i used it just nowyou can't, it's not comfyui compatible
>>109474084Alien videos that could fool boomers
>>109474107your mom's not comfyui compatible but yknow what i made it work.
>>109474113>>109474113
>>109474099a man of culture >>6208208
>>109472741
>>109474084black_widow_hulk.gif
>>109474112
>>109474118Kek
>>109474118kek
>>109474107?
>>109474118that's me
>>109474133now try the same gen without the lora, same steps and seed.
>>109474133the format of the lora is not comfyui compatible, look at your cmd console you'll get layers because the names of the layers aren't the same of what comfyui is expecting, this lora needs to be converted
>>109474138I did, and it worked but the quality was asscan't post cuz it's pron
>>109474084tailjob
>>109473988I don't really know how it works but i imagine its distilling the model using many small clips of many concepts and if they concepts happen to be in your prompt the model can then shortcut based on what is inside the lora. Someone correct me if I'm wrong. The side effect is the model have less flexibility.