Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109469842https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
blsd trd o frdshp
https://litter.catbox.moe/3hjtai.webm
>give h3 references of items>it still changes one single item and it ruins the gen completelyHow do I solve this?
Absolute cinemahttps://files.catbox.moe/fbp7h8.mp4>>>/wsg/6208736
>>109471509>How do I solve this?how many MP you're going for?
https://files.catbox.moe/2218u5.webm
>>109471509try sufficiently high res and the subject / retention_analysis prompt sectionsalso maybe like in older models using rembg or sam first helps to make the items even easier, though I haven't extensively tested if/how much h3 needs it.
>>109471509Prompt issue. Be more specific about your subjects. Use the correct prompt format
>>109471516I love how that guy is just quietly accepting his fate
>>109471525>>109471547The thing is that the higher res I go, the worse it gets. At 1mp it fails 9/10 times, at 2mp 100% of the time as it also changed the image dramatically.>>109471550I suppose I'll expand further on the prompt, trying it tomorrow.
>mfw Resource news08/05/2026>Inline Studio v1.2.62 - Minimax H3 Lora training still onlyhttps://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUFhttps://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3https://huggingface.co/Kijai/MiniMax-H3-TAE>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inferencehttps://github.com/6somehow/DAC-SPADE>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generationhttps://github.com/yizzz927/CAPE-T2V>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusionhttps://github.com/jd-opensource/JoyAI-Video-Edit>ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMshttps://github.com/YangYangGirl/ParVL>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diethttps://huggingface.co/JamesZar/OliveGemma-3B08/04/2026>stable-diffusion.cpp adds support for MiniMax-H3https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES timehttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editinghttps://github.com/IntMeGroup/MIEScore>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videoshttps://rathgrith.github.io/PeCA>Kandinsky WM 1.0: A family of models for Physical AIhttps://github.com/kandinskylab/kandinsky-wm08/03/2026>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>109471527STOP THIS
Is trooney here? I have a troonyou
>added reference audio to my prompt>told it to use the ref audio as a guide for the quote>not only did it work perfectly but it decided to give me the middle finger with my prompt and blend the concept of sucking a popsicle AND crunching down on it all within the 8 second time limitholy fuck i love this model so much. AND adding audio reference didn't slow the gen speed down, this finished in 17:55 this time.Anyway, this one's dedicated to you popsiclesucking anon.https://files.catbox.moe/iepr7x.mp4
>mfw Research news08/05/2026>SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrievalhttps://arxiv.org/abs/2608.03120>HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Planehttps://arxiv.org/abs/2608.03422>DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformershttps://arxiv.org/abs/2608.03082>Can T2I Models Draw from the Right Frame of Reference?https://arxiv.org/abs/2608.03357>Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editinghttps://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0>Self-Supervised Representation-Guided Generative Dataset Distillationhttps://arxiv.org/abs/2608.03218>Latent Reward Registers for Diffusion Preference Alignmenthttps://arxiv.org/abs/2608.03929>UniWorld-Design: From Pixel Generation to Layer-Native Designhttps://arxiv.org/abs/2608.03971>MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Bindinghttps://arxiv.org/abs/2608.03708>RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editinghttps://arxiv.org/abs/2608.03059>Efficient Video Dataset Distillation via Cluster-Guided Prototype Blendinghttps://arxiv.org/abs/2608.03269>Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfoldshttps://arxiv.org/abs/2608.03135>TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Modelshttps://arxiv.org/abs/2608.03057>Adaptive Two-Stage Visual Token Pruning for Efficient Inference in VLMshttps://arxiv.org/abs/2608.03112>Enhancing VLM Reward Models Through Structure-Aware Fine-Tuninghttps://arxiv.org/abs/2608.03875>Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understandinghttps://qwen-3d.github.io>When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardwarehttps://arxiv.org/abs/2608.03649
>>109471540>fucking everybody has a vpn to connect to the global internet. It's more of a discouragement than anything else.there's a difference between a chinese guy using a VPN and not saying that he's chinese so he's hard to spot, to a Chinese company like Alibaba using twitter, the CPP knows they're here, they allow it, that's how it works in China, the big companies can break the rules as long as they fuck the US in the ass, in a war there is no rules
stupid question maybe but can I generate minimax H3 videos without sound?and would this generate faster than with sound?also how you guys remove the audio so you can post it as a webm on boards that dont allow audio streams?
>>109471568now there's a name I haven't heard in some time. did he kill himself?
>res_multi / beta,>0.7 = no pattern artifacts just a bit of smearing>1.0 = insane pattern artifacts looks like a 0.4 genwhat causes this?
>>109471443wow its like i rendered native 1080p cool
>>109471558did you try to describe them as subjects and/or demand they do not get "interpreted" overly much via retention_analysis?
>>109471596i tried beta it sucked ass
pouring one out for rocketgirl anon who surely would have done great things with H3. RIP
>>109471527yuri is the only thing better than 1girl, because 2girls > 1girl, basic maths!!
>>109471516its huge IP knowledge is really what sells it to me, your only limitation is your creativity now
Anyone have any tips for when using the ref2va model with 2 audio inputs.I have:><Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 1> without copying the original signal.><Audio 2>: reference - <Audio 2> is used to provide a sample for the style of background music throughout the video without coping the original signal.But the video starts with <Subject 1> speaking and I dont get any background music until <Subject 1> stops speaking.Im just trying to understand if im doing anything wrong?>detailed description:>[Shot 1] The scene opens exactly on <Picture 1>. Music is playing throughout the video in the background in the style, beat, and rhythm of <Audio 2>.
reference prompt from grok to get the specifics just right (I added the music after):Cut 1 (0–2s): Dusty western street outside a classic saloon in the style of JoJo’s Bizarre Adventure: Steel Ball Run. <Subject 1> and <Subject 2> walk toward each other under the bright sun, both in period western attire with dramatic JoJo posing.Cut 2 (2–4s): Close-up on <Subject 1> as he winds up and throws the glowing green steel ball with intense focus, the ball spinning rapidly as it leaves his hand.Cut 3 (4–7s): The green steel ball flies in a perfect spinning arc and slams into <Subject 2>’s cheek with a heavy thud, stopping dead against his face while continuing to spin rapidly. <Subject 2> freezes in shock from the impact.Cut 4 (7–10s): Close-up on the spinning green steel ball still pressed against <Subject 2>’s cheek. <Subject 1> stands calmly and declares, “The power of the spin has beaten you.” Freeze-frame on the dramatic moment.pic related is source 1.https://files.catbox.moe/w508z3.mp4
>>109471634>it even knows the songkek, best model ever
>>109471642nah that I added, its not THAT good (yet)
>>109471642bro read
>>109471649>>109471652
So, if I wanted to talk to my waifu thanks to Minimax 3 how would I do it, if I didn't mind the lag:Input my question in the workflow make an llm answer and pipe the response towards Minimax along with a reference image and sample voice? Would that work?
>>109471632There's an entire section for background music.
>>109471595no, still alive and shitting up /e/
https://files.catbox.moe/3u8ym8.mp4video reference to video is a sin
holy shit, the second iteration of the grok jojo prompt worked even better.Cut 1 (0–2s): Dusty western street outside a classic saloon in the style of JoJo’s Bizarre Adventure: Steel Ball Run. <Subject 1> and <Subject 2> walk toward each other under the bright sun, both in period western attire with dramatic JoJo posing.Cut 2 (2–4s): Close-up on <Subject 1> as he winds up and throws the glowing green steel ball with intense focus, the ball spinning rapidly as it leaves his hand.Cut 3 (4–7s): The green steel ball flies in a perfect spinning arc and slams into <Subject 2>’s cheek with a heavy thud, stopping dead against his face while continuing to spin rapidly. <Subject 2> freezes in shock from the impact.Cut 4 (7–10s): Close-up on the spinning green steel ball still pressed against <Subject 2>’s cheek. <Subject 1> stands calmly and declares, “The power of the spin has beaten you.” Freeze-frame on the dramatic moment.https://files.catbox.moe/f3kjg8.mp4
>>109471669that generally should work. but first try without the llm always in the loop and simply do a test prompt with the waifu image and voice referenceand figure out if you need any of the features from VIDEO_PROMPT_WRITING_GUIDE_ref_en.md to make the audio or image reference(s) work like you want
>>109471694
>>109471694he ate krabby patties again i see
So all the loras for Minimax are shit.Will we ever get anime-focused nsfw loras? Or will it be like wan where it's all 3D shit and we have to constantly mess with weights. I'm also reading in reviews that configuring lora weights for h3 is a nightmare.
>>109471694THICC
>>109471694the funny part is that he would probably feel happier working as a prostitute than still working at Krusty Krab kek
>>109471712bro it's been like 3 days. chill.
>>109471712it just came out, daddy chill
>>109471727>>109471729Hey illiterate moron, did you misunderstand the question?Wan2.2 has been out for over a year and it barely has a single anime-focused lora. Civitai coomers are heavily biased towards 3dpd.
>>109471712be the change you want to see in the world
>>109471749>autist STILL repeating himselfForever a mentally ill freak
>>109471712https://youtu.be/TRTkCHE1sS4?si=pHwxwWd2kxtpnnfLFor real talk through enough informed compute at a problem and it will be solved, talking from experience and also solving heavier problems myself
>>109471699>copied gyros anime style from a stillsora couldnt do this
>>109471745>Will we ever get anime-focused nsfw loras?>just came outno, the answer is never. it will never happen.
>>109469731 >>1094698100.3 mp is just the spot hits the kino right at home; UHD sein just doesnt hit the same
>>109471761>only one guy in the world use this expressionsounds like the mentally ill is you
H3 quality is ass on a 24gb, I get better results with wan22
>>109471793but you can do ball park the same resolution and quality on a 5s or whatever clip there as on wan
>>109471745It's been a year and you have no dataset. You had a year to get a dataset, caption it, and work a part time McDonald's job for a week and train one yourself. And you didn't.
>>109471803hmm that's a long post... easier to complain...
>>109471793>H3 quality is ass on a 24gbwhat do you mean? you can go with high resolution at 24gb, yeah it won't fit the vram card but that's what offloading is for
>hotspot peak temp is 105C>30C delta between core and hotspot tempsI don't feel so good bros... I need to repaste...
>>109471813I promise they exist, people just don't like freeloaders.
>>109471289yes, anon. Now we all come to the cold, dark realization that we were never really as creative as we though we were and we just needed the right tools to express it...
>>109471793what are you talking about? wan 2.2 is limited to 1mp
>>109471570>popsiclesuckingkino>https://files.catbox.moe/iepr7x.mp4slop
>>109471829nope
Can h3 put subtitles along the bottom of the frame of your genned video that match the dialogue?
>>109471629and what a limitation it is
>>109471751this is actually a good benchmark for world understanding, once miku shatters into pieces then models actually understand physics and thermodynamics somewhat
>>109471587>can I generate minimax H3 videos without sound?I was just wondering the same thing. I don't particularly need audio for testing or for something silly I'll just be uploading here on 4chan.>>109471632You need to bind the audio samples to their references and make sure the task type is correct.
>>109471816I am going to render the same from both, will take 15 min.
>>109471820I'm very anal about my GPU temps, it reaches 75c and hotspot 90c when genning, put an external fan in front of the case and lowered them to 70c/82c
>>109471820learn the joys of underclocking
>>109471825gooners are kind of lucky in that regard.
>>109471843Sir please redeem the Beta scheduler so your gens will stop looking like blurry dogshite
>>109471838yeah, and sometimes it'll even do it if you ask for it
https://files.catbox.moe/ctwf8s.mp4
What is this minimax NSFW hype? I tried it and I just get shitty ouputs that merely resemble porn with shitty unnatural movement and weird looking genitals. I am doing something wrong or is it overhyped?
>>109471825>we were never really as creative as we though we were and we just needed the right tools to express it.unironically this, I was unconsiously nerfing myself when playing bad models because I was like "nah, too complex, it won't work", now I just can let it go and ask minimax to do whatever complex shit I want and it just workshttps://files.catbox.moe/1v0d0k.mp4
>>109471871i'm on the gigachad workflow, alpha sampler with a custom sigma.
>>109471843I put into prompt:>george hits the frozen iceblock miku with a baseball bat. She shatters loudly into bits that fly everywhere and rain down the fight scene.but I guess the model wants to maintain her there as a subject
>>109471883Use the reference model and swap in your favorite hag into a video of your choosing. It works perfectly.
never imaged we would get a seedance 2 level model localhttps://litter.catbox.moe/x6i1j9z62j7llj6z.mp4https://litter.catbox.moe/pgf6n624bl4dylns.mp4https://litter.catbox.moe/l4gexzs99rhfqnhr.mp4https://litter.catbox.moe/i1q53a87udb9np37.mp4https://litter.catbox.moe/fqfo94ysyc1716ap.mp4
>>109471883h3 treats anything going into the mouth as eating, no matter how you prompt it. they clearly didn't caption blowjobs or sucking properly. i'm hoping finetunes or loras can fix that.for now using a video reference is the best solution, however, its slow and doesn't work half the time
Swear my ComfyUI is borked. I was able to generate using H3 earlier now every time I load it just crashes. Da fuck is going on.
>>109471889I don't think so, champ. The model can produce clean, crisp videos even at low resolutions. You're doing it wrong. You may as well be using Wan 2.1
>>109471587output to mp4. You can upload that here. Also, you can just not include the sound noodle. I don't think there's a way to disable sound generation to make it faster.
>>109471843https://files.catbox.moe/nslbfo.mp4
>>109471922are you the guy vibing his sage install?
>>109471915>seedance 2it's not seedance 2's level, let alone seedance 2.5 level, but it's really good nonetheless
Fellas I sold a kidney and have some upgrades coming in. What OS is everyone running these things on? You guys just on Win 11? Last thing I want is to run these models on a spyware OS.
>>109471896Thanks I will try that
lol I taught the reference model what a stand is. need to tweak but...yeah.<Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>. <Subject 3> is the character in <Picture 3>Cut 1 (0–3s): Dusty town street in the style of JoJo’s Bizarre Adventure: Stardust Crusaders. <Subject 1> and <Subject 2> stand facing each other outdoors under bright daylight. <Subject 1> adjusts his hat. <Subject 3> suddenly materializes beside <Subject 1> in a flash of energy while <Subject 1> himself remains still.Cut 2 (3–7s): <Subject 3> unleashes a rapid “Ora Ora Ora Ora!” barrage, its fists delivering a flurry of powerful punches that hammer into <Subject 2> repeatedly with heavy impact effects and speed lines. <Subject 1> stands calmly in place, only slightly moving his hand as if controlling the Stand.Cut 3 (7–10s): The barrage ends and Star Platinum fades away. <Subject 1> lowers his hand, looks away, and calmly says, “Yare yare daze.” Freeze-frame on <Subject 1> in his classic pose while <Subject 2> is left reeling.https://files.catbox.moe/1amgv6.mp4
>>109471942it legit is, all we need is the 2k upscaler they promised
>>109471915>https://litter.catbox.moe/l4gexzs99rhfqnhr.mp4wtf are you doing Jonny??
>>109471950Make sure to check the prompt guide. I posted a working prompt in the last thread if you wanna peep it and modify it to suit your tastes.
>>109471922That's why I don't install custom nodes.
>>109471962link it
>>109471938>>109471964Nope. I just used a basic template and added a clean vram before the vae decode. im tech illiterate.
>>109471956it's not, even in your video there's a lot of morphings during the fights, seedance doesn't have that issue, and an unpscaler isn't gonna fix that
Fuck how do I prevent characters from talking in minimax? I want vocalizations without dialogue.
>>109471859I don't think that's gonna lower my hotspot temps by 20C unless I go like 60% power limit.
>>109471980a latent upscale is exactly what would fix that cheaply. Natively doing it would work as well but that would take way too long
Man if only Minimax can lower FPS natively16 fps will save generation time
>>109471928>output to mp4. You can upload that here.does it remove the audio automatically?
>>109472004yes
>>109472004You just need 2 create video nodes, one with the audio the other without. it's not complicated.
>>109471989>a latent upscale is exactly what would fix that cheaply.really? do you have any examples of that, now that's interesting
>>109472020minimax already said they plan to release the module to do so. It uses the model itself to upscale
The fuck am i doing in these last 2 days80% of my gen is shit. Fuck Minimax
>>109471922try the --disable-pinned-memory flag
>>109471975nigger>>109470669
>>109471984Anime character ? Good luck lol because they trained from DEEN Dragon Ball Super animation lmao
>>1094717760.35 is the ultimate megapixel setting
>>109471994>16 fps will save generation timeKrea2 can do 1fps 1 second videos. Try that
>>109472031skill issue
>>109472038Yes anime. It doesn't happen with all gens, but this one particular one is a nuisance.
>>109472037also helpful for voices:<Audio 1> is the voice reference for <Subject 1>, containing a spoken English vocal layer.
is there any way to snap these stupid little things to the top free space up there? i can undock the run button but that's not really solving the issueshit always randomly blocks a node its so annoying
>>109472050Wait until 2d 3d lora appears. I bet they overtrained this
>>109472015do I need custom nodes for that?>>109472013ok I will try
Rachel Blevins refuses to answer my DMs and she WILL submit to Islam for her sins.https://files.catbox.moe/cy6o87.mp4
16min for 10sec 1mp. hmmmmm......virtually no artifact tho.
>>109472075geopolitics huh
grok pretty good for idea promptingCut 1 (0–2s): Soft stage lights glow. <Subject 1> stands center stage gently holding her electric guitar, looking a little shy but focused. <Subject 2> sits behind a drum kit in the background. They begin playing a softer rock melody.Cut 2 (2–5s): Music-video style cuts. Close-up of <Subject 1>’s fingers moving smoothly across the guitar frets, then a gentle cut to <Subject 2> keeping a steady, relaxed drum rhythm. Warm stage lights shift softly with the music.Cut 3 (5–8s): Wider shot of both performing together. <Subject 1> plays the guitar with calm concentration, only slight natural movement, while <Subject 2> supports her with smooth drumming. Soft intercuts between their faces and instruments.Cut 4 (8–10s): Final gentle shot as the melody continues. <Subject 1> strums the last chords quietly while <Subject 2> eases off the drums. Freeze-frame on the two of them playing under the soft stage lights.https://files.catbox.moe/n9d795.mp4
>>109472090Upload porn and they will reject it. Grok is dead now
>>109472060Yeah, audio is an after-thought for me but I'll have to add that.
>>109472097well, still good for ideas, same with google ai searchif you want to prompt something extensive it saves a lot of time, basic stuff you can just type yourself.
How long until we're out of the video slop epidemic
>>109472100just install ollama and set up a good system prompt, it will write whatever you want in whatever format you want.
>>109472112sorry to break it to you but, probably never.
>>109472097>Grok is dead nowwhy do I have a feeling Elon will bring back some Grok coom kino because he can say that Minimax has opened the pandor's box
>>109471855Guess what's what.Wan or H3.https://litter.catbox.moe/clahzn41c5kdlpaw.mp4https://litter.catbox.moe/hhmpqoem750pfo6u.mp4
>>109472112>video slop*peak
>>109472112video is more dynamic, and you can still use krea 2 and klein edit to make your source material to make videos with too.
>>109472120I think it's simply too big of a legal liability and he's target number one.
>>109472112video slop is inherent to AI because it allows anyone to make a video. It's always going to be slop no matter how good the models are.
>>109472124>pedo and pedo.
>>109472124the bottom one is wan because its known for making the subject have a seizure unprompted
Altina
has local diffusion improved much in the past year? i found it meh when i tried it earlier this year, wondering if it's worth taking a dive again
>>109472148Do Tio instead. Altina sucks.
>>109472160nah just go back to sleep unc
>>109472136Yes, that's the entire point. I was not asking how long until AI videos get good but how long until the thread stops being 95% bad AI videosI could at least enjoy a random image if I didn't go looking for imperfections in it but every single video is disgusting
>>109472160>has local diffusion improved much in the past year?let's see, 8/5/2025 we didn't have Z-image turbo, Klein, LTX, Krea 2 and Minimax, so I'd say we improved a lot yeah
>>109472160We're back to levels never dreamed possible. All fields, LLM, image, video.
>>109471844does telling it something like "no dialog or music in the scene" not work? I've heard that works in the seedance or one of the sass models
>>109472142I call this "dynamics" but you're right.The motion on H3 is ass, physics as well (ankles clip).Equivalent rendering time and upscale.
>>109472169>the thread stops being 95% bad AI videoscan't happen, that's the strugeon law, if every video is super, no one would behttps://en.wikipedia.org/wiki/Sturgeon%27s_law
I'm trying out H3 on stable-diffusion.cpp but am violently OOMing on 24GB VRAM + 32GB DRAM.Before I try q2 instead if q4, which components of the model need to stay loaded simultaneously at any one time? And has anyone else tried this method?
>>109472170i dont know any of these so sounds like it's worth trying againlast time i tried i think chroma was what everyone was using
>>109472185sounds about right.
>>109472184>The physics on H3 is assnonsense, look how those boobs jiggle!https://files.catbox.moe/scgf72.mp4
>>109472200>I'm trying out H3 on stable-diffusion.cpp but am violently OOMing on 24GB VRAM + 32GB DRAM.wtf are you doing, use comfyui it lets you use int8 convrot at 24gb of vram with fast offloading shit
>>109472211lel.I've seen good physics from H3, but none on my setup.
>>109472142Wan also has the dust flying around on those types of dark backgrounds - clearly this was with an image with the background removed
>>109472160local has caught up to, and in some areas even surpassed api. it's surprising, but at the same time not really, because the corpos mission is to make profit, not the best AI possible. they cut corners and mostly offer convenience and speed thanks to their data centers
>>109472148Altina LTX
comfy confirmed the official API WF is 50 step 720P upscaled to 2K
>>109472200it isn't respecting the max vram limit or offloading properly. give leejet a break. he doesn't get weights early
>>109472220we are trying to leave the cumfart griftosystem goy
>>109472241>0.92 MPdoable>50 stepsoh hell nah man, I'm waiting for the turbo loras
>>109472238any recommended image models for degen shit that isn't safetymaxxed
>>109472252krea 2 with nsfw loras is the current meta
Anyone have a good prompt for subtitling the dialogue on the bottom of the screen? Trying to make it work with the official prompt template but can't get it working.
>>109472267Just use a video editing software?
I'm going to bake the will smith spaghetti again and add the ranfaggot rentry next time
>>109472249comfyui has functionality that doesn't exist in those shit a1111 clones.
>>109472220I like the commandline and am greatly annoyed by venvs in general and by drowning in nodes in Comfy.I use stable-diffusion.cpp for imagegen, knowing it's slower and less resource-efficient than ComfyUI, but it was worth it for me. Seems to have hit its limit now, however.
oh, I see now...https://files.catbox.moe/6vpgqo.mp4
>>109471984tell your clanker to prompt it
>>109472277so does ggml. llm bros are living further down the tech tree and don't have to deal with python niggardry. comfyui is past it's prime already and can be rewritten and relicensed with ggml. that is the point of it. no more fucking retards asking how python works or shooting themselves in the foot
>>109472265does the base suck? i always hated having to use loras since they tend to be narrow
bocchi and the goose playing guitarCut 1 (0–2s): Bright stage lights flash. Natural handheld camera captures <Subject 1> standing center stage gripping her electric guitar with focused energy. <Subject 2> stands beside her also holding an electric guitar. They launch into an uptempo rock song in the style of the Bocchi the Rock opening.Cut 2 (2–6s): The camera moves with natural, slight handheld motion around <Subject 1> as she plays fast, energetic guitar riffs with growing confidence. It drifts smoothly to also show <Subject 2> matching her energy with his own rapid guitar playing under flashing stage lights that pulse with the music.Cut 3 (6–10s): Natural wider shot of both performing at full energy. <Subject 1> and <Subject 2> both strum and pick rapidly on their guitars. Freeze-frame on the two of them mid-performance under the bright stage lights.https://files.catbox.moe/uktx9w.mp4
>>109472314insane how good the camera movements are, try to do that with wan or ltx it would end up as some smeared garbage
>>109472267Ref model[Shot 1] Subtitles on bottom of the screen: "lah lah lah".[Shot 2] Subtitles on bottom of the screen: "nah nah nah".
>>109472300Anon, you're an idiot. An LM doesn't magically know how to get around issues with the model.
Can you have 2 subjects in 1 input image?
>>109472296Using a reference video to get the movements?
>>109472267......and say with subtitles shown "Lloyd san, i know you have Elle... but would you having an affair with me?"
>>109472301>comfyui is past it's prime alreadycomfyui is literally the gold standard. no one gives a fuck about whatever slop shit you think would be better, probably rust because you sound like a tranny. comfy works extremely well and is insanely flexible, that is the bottom line.
>>109472351>gold standardgold no, standard yes
>>109472339yeah. just say s1 is the character on the left and s2 is the character on the right
is video gen still like 5 minutes for 5 seconds on consumer hardware or has it evolved
>>109471527tsundere X tsundere yuri is the peak of what humanity has to offer>>>/wsg/6208791https://files.catbox.moe/439aki.mp4
>>109472362Big if true
Chaining transition shots together infinitely long is an option. https://streamable.com/qavxej
>>109472314what's the source of that music? I like it
>>109472358what is the alternative?in one corner you have comfyui, a fucking kino factory that does anything and everything you could think of.and in the other corner.. some gay prompt window uis, clones of abandonware from 2022. but hey, they were written in c++ or whatever, because that matters.
>>109472351>comfy works extremely well and is insanely flexibleerm, no
sol attention + sage + UC_MiniMaxH3Cache from https://github.com/silveroxides/ComfyUI-UtilsCollection = fast as fuck boy.
wait did civitai do a purge of nsfw stuff? which aggregators do people use now
>>109472398Never said there was an alternative but just because comfy is the best we have doesn't make it great de facto
i just spent 6 hours straight genning with h3. am i addicted?
>>109472400>sol attention>only works on sm90/100 3xxx bros...
>>109472411nah its addictive as fuck. Reference video's potential is endless
>>109472301Sadly hasn't trickled down to diffusion yet. stable-diffusion.cpp is the only one in the field I know and, unless I missed some commit notes, lacks a lot of the options that ComfyUI has, like this new sol attention thing or convrot (which AFAIK is an alternative to gguf however so makes sense).
>>109472413it's an extremely popular model, there'll be stuff for us
>>109472411>am i addicted?I won't call an addiction, because it doesn't hurt you, and I'm like you anon, I'm glad I'm feeling some genuine joy over a hobby, it's been a while that didn't happen
>>109472435>because it doesn't hurt youit fries your neurons anon
what's the meta for anime locally? anima still? or krea2
>>109471702I did it. Gemma-31b powered realistic waifu can speak to me.https://files.catbox.moe/ofla5c.mp4
>>109472444talking to you fries my neurons
>>109472444>>it fries your neurons anonsleep fully restores the neurons
>>109472405it's the best we have because it's the best ui for this shit, it also has a massive community of people developing nodes for it and full industry support.
>>109472321thank you very much saar
>>10947238>infinitely long What is your process, king?
>>109472389randomly generated in the clip.
this is the best setting for https://github.com/silveroxides/ComfyUI-UtilsCollection btw. It has no degration at all even with fast movements but is still a 40% speed up
>>109472466>>109472386Oh, fuck I meant to quote you.
>>109472471really? it sounds so good I thought it was a real music, Minimax is a beast
>>109472473is this basically the same thing as that chink node with the same name?
>>109472486yeah but that chink node OOMs my card when I do a second gen somehow
>>109472486no, the chink node hurt quality a lot even at low settings
>>109471871>Sir please redeem the Beta scheduler so your gens will stop looking like blurry dogshiteSorry anon I queued up 20 gens and stepped away before I read this so we'll have like 30 shit gens to get through first in like 2 threads from now
>>109472473I guess I have to remove spectrum for this right?
>>109472495Thanks I'll switch.
MiniMax is like fucking crack cocaine lord help me
>>109472504From my almost full day of testing I can tell you so far the best gen-time/quality cope node is that one.
>>109472473holy fucking bloat repo. not installing
5 MP test, 3040x1768https://files.catbox.moe/p84fby.mp4
>>109472511>MiniMax is like fucking crack cocaine lord help meEnjoy it to the max anon because it will never hit this hard in a month
>>109472511It's gonna look like shit to you in two weeks.
It even emulates the hair clipping in video games lol
>>109472514>From my almost full day of testing I can tell you so far the best gen-time/quality cope node is that oneSomeone point an AI at all the cope nodes and get them to summarize the differences for a layman
>>109472522>It's gonna look like shit to you in two weeks.nah, unless we got seedance 2.0 at home, maybe BFL will surprise us (kek)
>>109472527spectrum
I'll have to use "that"
>>109472531Censored vs. uncensored, I don't know.
>>109470393@0.7mp lowk wasted af with a gen that starts that close tbhdesu
>>109472473im only using KJ nodes. anything else is guaranteed snake oil
why DID minimax make an open weight video gen model? It is because they know they cant directly compete with kling/sora/seedance?
It has been months since this thread was this active. MinMax is blessed model of frenship desu
>>109472518I'd be pleased with a 1080 upscale to 2048.
>>109472473I installed it but i don't see this node?
>>109472566I really don't know, they could have survived in the API space, it's 3x cheaper than seedance and it's perfect to make simple scenes
>>109472566they wanted to see me make porn, low key
>>109472580weird, because I just git cloned it and I found the node
>>109472473What the node setup for this, Sage into shift into cache?
>>109472566they threw the gooners a bone after alibaba abandoned us
kekhttps://files.catbox.moe/imh9fp.mp4>>>/wsg/6208799
>>109472566>why DID minimax make an open weight video gen model?china's goal with ai is dominance in both closed source and open source. so for they are winning
>>109472586>>109472473why on earth would I want to download some dudes collection of other peoples nodes? i'd rather just download the nodes I want so I can keep them up to date.
Is he right?>>>/a/289931840>this kind of generative AI is like a slot machine, you pull 3-4 times in hopes of getting a few seconds of usable stuff, but there's no consistency and the amount of time, editing, and tokens you'll spend trying to get anything useful is going to cost more than just hiring some SEA or SKorean animation studio farm.
You can use h3's text encoder as an LLM. Can anyone abliterate it and fine-tune a lora with h3 optimimized structured output?Then we could easily prompt up-sample using the encoder and lora in the same workflow and encode normally afterward.That would be a pretty neat quality-of-life feature.Given the level of uncensoring it will remain our base for a while.
>>109472605>other peoples nodes>>109472495
>>109472602*so far
>>109472610>abliterate italready done
>>109471751>miku trooned outsad news
>>109472611so what, you put em through gpt sol/claude and fixed them?
>>109472566To cause legal troubles in the West with AI models, probably.
How long until someone makes a full length film/season with this?
>>109472518another 5MP testhttps://files.catbox.moe/78ce5a.mp4
>>109472632the chinese have already done this several times over
>>109472605its not other peoples nodes. That is silveroxides
>>109472586Ah, yeah that was it. I installed through Comfy extension and that version didn't have the nodes.
>>109472637Gen time/specs?>>109472638Source?
Are the pruned models okay or are the noticeably lobotomised?
>>109472637>5MP test>blow it on animated shit that can't even showcase it
>>109472518it's literally real life
>>109472566because api is dying and it makes sense to build a userbase that is actually going to use your model.
>>109472648>Source?i've seen several slop series produced that make their way onto youtube. specifically with seedance2. too early for it to have been done with H3 yet
>>109472650they are more than ok, it has a similar quality to the original model, the original model's architecture is outdated, and it could be optimized by replacing 13b of layers with a table or something
>>109472660They're a company. What's the point of a userbase consisting only of demanding leeches?
Anyone else getting dialogue completely destroyed when using EasyCache or MiniMaxH3 cache? Even without them it's really echo-y. This is at 1 MP with both disabled. If I enable either it's basically indecipherable garbage audiohttps://files.catbox.moe/74219e.mp4
>>109472401No they didn't, even the Minimax stuff is there
>>109472448nice. i think you can still improve on the sound. e.g. preprocessing your sample with AudioSR or some other speech restoration / another tts' voice cloning type thing or some other thing so the odd noise maybe goes away
has anyone tried extending the middle of a video from a single video ref? so that the beginning and end stay the same? or would you have to cut that up into two video refs?
>>109472605I used chatgpt to review the repo then isolate the minimax h3 cache node into a separate downloadable archive
>>109472607Yes he is right APIs suck and are extremely expensive because you throw away a large percentage of your tokens. Local changes a lot, it allows unconstrained experimentation and you'll soon be getting better results.
Hello fellow Vaisyas. Please do the upload of workflow for MinMax sir thank you.
>>109472673>Anyone else getting dialogue completely destroyed when using EasyCache or MiniMaxH3 cache?yes, that's why I use spectrum insteadhttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
>>109472682prompt?
>4 sec gen took 70 sec>8 sec gen took 200 sec>10 sec gen took 450 sec.........This must be a bug, right ??
>>109472690just use the official comfyui template, it works fine
okay this is how you get a character to play music from a source input (ref audio 0) apparently:<Subject 1> is the anime girl in <Picture 1>.REFERENCES:Use <Picture 1> for the background stage setting. Use <Audio 1> exclusively as the instrumental guitar performance track.STAGE PLAYBACK RULES:The audio output of the video must feature the sound from <Audio 1> completely isolated. Mute all procedurally generated AI background noise, voices, and ambient sound effects. The final video must contain zero audio artifacts except for the exact, unaltered guitar music file provided in <Audio 1>.SHOT & SYNCHRONIZATION:A medium shot of <Subject 1> playing the guitar on the stage. The strumming and hand movements of <Subject 1> must perfectly synchronize with the physical tempo, rhythm, and notes of the guitar track in <Audio 1>.
>>109472566china dominates the hardware manufacturing side, america the software side.how china wins: make software worthless.easy.
>>109472693For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.integrated_multimodal_description: SHOT 1: 2D-animated, the camera holds a static shot. The girl from <Picture 1> uncrosses her legs then crosses them the other way. Then she looks down on the camera, blushing slightly. She is in constant subtle motion. The girl is breathing softly for the entire video. The girl is wearing black shoes.overall_soundscape: crossing legs on a couch, indoor ambience,non_diegetic_music: N/A
>>109472673i'm having a few audio issue too and I also thought some of them came from EasyCache but have not REALLY confirmed it, i'm not that confident in the quality of my samplesif >>10972692 doesn't fix it sufficiently maybe you can get further improvements via audiosr or comparable on either the input audio or just filter the output with it before the video is merged with the audio? i also got rid of a lot of shit effects of default LTX and some TTS models that way.
>>109472696Unfortunately, the time dramatically increases the gin time. I have a 5090 and going from 10 second to15 seconds changes the gen time from 15 minutes to 2 hours.
>>109472671in creative spaces having people prefer your product or ecosystems pays dividends in the long run. for decades adobe didn't give a fuck about piracy, they dominated the market and made a fortune. i don't use any apis, i would pay money for krea because i understand the model now.
>>109472713ok, if it goes that high it means your overfilling vram and using CPU. You can use FFN chunking from kijai's nodes instead
>>109472713It must be bug. Even LTX and WAN didnt do this
>>109472717>creative spacesbroski... the artists and creators fucking loathe vidgen
>>109472712AudioSR demo https://audioldm.github.io/audiosr/ - it is *nearly* an universal just fix the ugly bits band-aid solution
>>109472722sure they do. unless it's seadance or one of the "good" models, right? that stuff is the real deal!
>>109472718>You can use FFN chunking from kijai's nodes insteadI have no clue what that is, how does it work?
hello sears i'm leaving this image here for no reason whatsoever do not do the needful, signed yours deemfully
>>109472651lolz, yeh anime could be in 240p and look just as good
>>109472675I wanted a girlie voice but the clip I used seems to be too muddied for this model, with TTS cloning it was good enough.I connected the image to gemma and asked to show me what she was wearing, the text came out fine, but the audio is garbled, I'll find a better reference for her voice.https://files.catbox.moe/8qrerc.mp4
I just wanna remind how far we've gone. This is H3 t2vhttps://files.catbox.moe/fv8ruj.mp4And this is LTX 2.3https://files.catbox.moe/2u8nuy.mp4
>>109472729you need to exit your little bubble, all AI gen is considered theft from artists. thus, they hate it
>>109472696update comfyui
>>109472748ghiblislop
>>109472741okay i wont
>>109472713comfyui 0.30.2 ? Any changelog ?
get me your most absurd gen idea. i do not care how much sense it does or does not make
most of the issues with extreme gen times comes from messed up or outdated comfyui setupsmy old portable was made during wan2.1 release and was all sorts of fucked when trying to run new stuffbetter to start over clean every once in a while
>>109472607Not really. Unless the average slot machine gives you ways to increase your odds.
>>109472758and you need to exit your bubble and look at what actual working artist are doing, not just transexuals on twitter mad that their gooner commissions dried up.i guarantee you at least 75% of concept artists in the game/film industry are using ai.>t-they said they aren't! they lied.
>>109472801i did this recently by manually updating the venv to the latest recommended shit and it upped my gen time by a lot while reducing vram usage lol
>>109472801This is why i backup my comfyui portable before downloading Minimax.
>>109472799Seconding this but instead make it a Family Guy bit. I haven't seen one for H3 yet.
>>109472823If I just delete my old one and make a new one will that mess upp all my custom nodes or will the comfy manager manage that on start?
>>109472820>75%tbqh if you said 25% id believe you, considering there's flagrantly pro-AI game dev companies like SHIFT UP emerging
>>109472518smells like beta scheduler.
>>109472799Katara milkbending Azula
>>109472799Bibi eating Donald trump, Frontal shot of Donald looking at the camera while traveling Bibi's intestine, Bibi shitting out Donald
Any easy way to get cuda 13 on stability matrix?
>>109472794I'm on 0.30.0. Updated yesterday. Has something changed?
https://files.catbox.moe/40r0bf.mp4
>>109472160its way better now for image and video generation.
>>109472837you realize that most concept artists kitbash right? they literally just google pictures of shit they want and chop them up and slop some paint on top. i don't really understand who you think is spending hundreds or thousands of dollars a month on these goated api models if not working artists in the industry.
>>109472866>hundreds or thousands of dollars a month on these goated api modelsIt's not the artists who chose to use them. It's the suits that get offered a new product that advertises "make slop to make money, but faster" so they buy in, and force everyone to use it becase muh "it's the future!"
>>109472845im listening but im not sure if we have the technology for this yet.>>109472849fuck off
Christ so many newfags that don't know what quadratic scaling is
>>109472862can i get a catbox for that gen? curious to see model + prompt used
>>109472586Alright, messed with it a bit. I get about 10% faster gens than with Spectrum.Cannot combine this and spectrum, gives me a node error. Probably need to fuck with settings, might be a gpu specific kinda thing idk. What GPU were you running for a 40% speed up? And what video settings? I'm running a 4070, doing 0.3MP 16:9 videos @ 10 seconds. ~450s with the cache node, ~490s with spectrum
>>109472877nibba talm bout quadtratic lol
>>109472875>im not sure if we have the technology for this yetLife can be so cruel
k2 won
>>109472883>Cannot combine this and spectrumit's one or the other, they are both meant to skip steps
>>109472883>Cannot combine this and spectrumCache does the same thing, but worse.
>>109472877nah it's actually linear, you're just dumb
kino bread
>>109472883https://files.catbox.moe/u6ou7n.mp4
>>109472888lmao
>>109472908WE MUST BE BETTER MENDO NOT STACK CACHE NODES
>>109472872>It's not the artists who chose to use them.except for all the artists that choose to use them, because in the real world results matter. no one gives a fuck about some shitter who thinks every brush stoke is a silent tear for their soul.>scan your art>train a lora>you just streamlined 3 hours of your daily work
>>109472881No
Why is no one talking about how to improve the quality, only about how to gen faster?
>>109472923unironically based response.
>>109472923BASED GATEKEEPIE
>>109472926we dont mind sacrificing some quality for more speed. if you want more quality then increase the steps and mp
Original OP seems a sleep. I wont take any chances >>109472936>>109472936
>>109472930(he's on civitai btw)
>>109472926For best quality you go high resolution and high steps, that's it. remove all the optimizations.
>>109472926you improve quality by generating at a higher size. you want to iterate quickly so you can tell if you have a seed that's worth the time it'll take
>>109472640I dropped both implementation in Claude and it said it basically was the same thing. just that silveroxide was cleaner. but functionally the same.
>>109472944keke
here you go anon, ignore the other faggot. https://files.catbox.moe/fmh1oy.png
>>109472926I want to make 10 gens in under 30 seconds
>>109472945throw me a bone dawg
>>109472923reminder that you can just train on gatekeeper's results in order to copy their style for free
>>109472980BONE: THROWN.https://civitai.red/user/obinna7713/posts
>>109472881>>109472967missed reply.>>109472988i wouldn't do that, op can easily achieve my style depending on the checkpoint he uses. Just avoid using the overly baked nsfw lewd krea2 checkpoints.
>>109472942I'm the opposite I don't mind waiting.>>109472946>>109472949Yeah, that’s what I’m doing. Unfortunately there’s a limit to how high I can push the resolution. I hope we get a tiled upscaler or something that can alleviate that.
>>109472967im assuming you trained a character lora for her, looks quite good. gonna have to try it myself
>>109472862
fl2va is cool but only if the camera is going to continue to remain the same distance from the subject the whole time, otherwise it just fucks up the face when it pans in with no other references
>>109473072what the FUCK no you can't DO THAT!
>>109472448 >>109472747https://github.com/diodiogod/TTS-Audio-Suite/blob/main/docs/VOCAL_REMOVAL_GUIDE.mdwould have tried myself but i have some issue with a python<->C lib for this. the node pack is probably anyhow going to be useful https://github.com/diodiogod/TTS-Audio-Suite/tree/main#features
Why are we baking threads at post limit? This board isn't fast enough for that.
>>109473012change the pg rating settings if your looking for lunafreya and amalia (wakfu) lora.>>109473047i indeed trained amalia (wakfu) lora with images from season 1 of the French cartoon. the captioning was autistic but i liked the results. Working on the captions for aranea highwing (final fantasy xv) lora but it will takes days. It's tough, challenging and effort focus process for me but others seem to do it piss easy with no sweat.
>>109473115nah i was talking about ilulu. that's cool though
>>109473105t2v prompt from social upvoting site https://litter.catbox.moe/4k2yrlbbov8lcdkw.txt
>>109473112because if we don't we'll get a schizo troll bake.
>>109473135all of them are so who cares anymore
>>109473196>t. the schizo troll in question
>>109473231>he's flooding his asshole with bacteria that will transform into worms and give you cancerngmi
help a new smoothbrain out here, when i use krea2 fp8 with comfyui just the base model, image generation is pretty quick at around 15s, using 24gb vram, and i see there's room in vram usagewhen i try to use a lora, vram is 99% taken, and image generation takes forever (minutes rather than seconds), not sure if this is expected? im guessing memory usage is overflowing and it's slowing things downim just using the basic krea2 example workflow from comfyui
>>109473072From oversized milkers to pure perfection. Good job.