Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109457662https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Zhttps://huggingface.co/Tongyi-MAI/Z-Image>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Wanhttps://github.com/Wan-Video/Wan2.2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
Does a high res fix pass help with Krea2?It seems kind of pointless when you can make images at such high resolutions out the box
I hate "ecelebs" especially men who offer nothing for my penis
>>109459116dunno if it helps visually but it could potentially make seed hunting & prompt iteration faster
Blessed thread of frenship
RTX or seed2VR ?
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.mdMan, ref2va is complex. I don't think we've had media generation anywhere near as complex as capable as this one but it's pretty challenging to get a full grasp of.
For anyone using r2v in mini, have you had any luck replacing likenesses in videos? I'm trying to use the syntax from the prompt guide but it's not clear to me why it isn't working. It seems inconsistent. If I want the motion and action from a video is that an "attribute transfer" or a "weak reference" or "partially preserved"? Or am I overthinking it?If it helps I'm only trying to use the reference video in one of the shots of the output.
>>109459151not sure but google ai search has been helpful for figuring out reference prompts and so on
r*ddit is going to get this shit bannedhttps://www.reddit.com/r/StableDiffusion/comments/1vf9wv7/breaking_gooner_minimax_h3/
the future: people gambling on the outcome of AI reference fighting videoshttps://files.catbox.moe/n3i2hz.mp4
>>109459168million dollar idea?
>>109459168>>109459173saltybet2.0?
https://x.com/ostrisai/status/2084642732610396411https://x.com/ostrisai/status/2084648469877141998ostris making both turbo 4 step lora AND undistillation lora
>>109459173>>109459180sure, if you prompt "one of the men falls down" AI has to guess who. cant rig it. just plug in character reference images.
>>109459168what the fuck thanks for the idea idiot that's just saltybet + ai and the only thing you need after that is to make it actually good>>109459180>saltybet2.0?imagine the insights degenerate gamblers will obtain for us about these models. they'll learn that hitler beats floyd 57% of the time or something like that
We can now prove so many thought experiments>We would win? 100 chickens vs 1 lion>100 chicken sized bears vs 1 bear sized chicken.Endless possibilities.
can you 4chan hackers stop ruining it for everyone? you're going to get this shit banned if you keep genning problematic stuff and posting it online
>>109459197we have been blessed with TWO models. two. and this reference model in particular can do all kinds of crazy shit. im simply using 2 image nodes. it can take audio AND video as input.
>>109459200boohoo nigga the model is out
>>109459197what happens if a planet of fire and a planet of ice collide
okay I need to tweak it, but the gen works: asmongold vs floyd.the black man Floyd in image 1 is having a fight with the man named Asmongold in image 2, in the streets of New York during the day. Asmongold is wearing a white tshirt and black track pants. Floyd and Asmongold have a fist fight and Asmongold says "you're getting hit by a chud, faggot!", after several hard punches to the face of Floyd, Asmongold knocks Floyd to the ground and Floyd says "I cant breathe!" as Asmongold stands over him.assign a name to reference image 1 or 2 then you can t2v prompt as usual.https://files.catbox.moe/qplluy.mp4
>>109459200Let them archive the models while we do what we must.THERE'S NO BREAKS ON THE GOONING TRAIN!
>>109459200From what I've seen reddit is a lot more likely to get it """Banned""" with all the copyrighted stuff they post.Nobody gives a fuck about what people do on here.
>use comfyui desktop because it just werks as far as installing python does>desktop version uses fucking chromium which takes up way more memory than it shouldis there any way to just use comfyui desktop while having it boot into my main browser instead of chromium?
>>109459218>Chinese model always makes black guy loseI'm noooticing.
>>109459227no lol
>mfw Resource news08/04/2026>stable-diffusion.cpp adds support for MiniMax-H3https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES timehttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editinghttps://github.com/IntMeGroup/MIEScore>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videoshttps://rathgrith.github.io/PeCA>Kandinsky WM 1.0: A family of models for Physical AIhttps://github.com/kandinskylab/kandinsky-wm08/03/2026>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md>Raylight 1.7.2 Adds MiniMax Support, 2x Speeduphttps://github.com/komikndr/raylight/releases/tag/1.7.2>Scaling Properties of Text Conditioning in Visual Generationhttps://heheyas.github.io/context-scaling>Retrieval-Driven Training-Free AI-Generated Video Attributionhttps://github.com/renxi-seu/Video_Attribution>A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Sampleshttps://github.com/zfu006/SSG>ComfyUI MiniMax H3 Image Studio (Experimental)https://github.com/astropuzzo/ComfyUI-MiniMax-H3-Image-Studio>MiniMax H3 — NVFP4 (Blackwell)https://huggingface.co/lilcheaty/MiniMax-H3-NVFP408/02/2026>MiniMax H3https://huggingface.co/MiniMaxAI/MiniMax-H3>MiniMax H3: Repackaged model files for ComfyUIhttps://huggingface.co/Comfy-Org/MiniMax-H3>MiniMax-H3-INT8-CONVROThttps://huggingface.co/Gluttony10/MiniMax-H3-INT8-CONVROT>MiniMax H3 COMMUNITY LICENSE AGREEMENThttps://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE>LoRA Dataset Studio: LoRA workflow in one tabhttps://github.com/perfectgf/lora-dataset-studio>comfyui-vram-trackerhttps://github.com/PuppetMasterAI/comfyui-vram-tracker
NOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO YOU CAN NO JUST MAKE GEN WITH MY OPEN SAUCE MODEL NOOOOOOOOOO
trying .9 megapixels, 5 second video on REF model with 5070ti with just a single 5 second video reference>219.33s/itFL model can do 1 megapixel 5 seconds just fine at under 20s/it. Why is it 10 times faster than REF? How do I into REF model at decent resolution? What slows it down?
>>109459159is that so
>>109459082I did a test running this with this flags, and still cannot gen 15 sec in 1mp I can only go for 9sec in that resolution... I don't know what is wrong. Someone with a 4090 and 64gb of ran which can confirm what is the max resolution possible and the time you can gen?
>>109459218we're not thinking with portals.it's not just bet on who wins, you bet like you're betting on roulette. yeah red or black is either of them winning, but you can bet on green (a tie, or some easter egg / funny resolution) as well, or on some division of the numbers (e.g. you can do a sidebet that george floyd pours a bottle of Hennessy on asmongold leading to victory which happens in 10% of fight videos)
>mfw Research news08/04/2026>Token Radius Attention for Efficient Video Generationhttps://arxiv.org/abs/2608.02504>CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generationhttps://hanxjing.github.io/CultureVidBench>EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generationhttps://arxiv.org/abs/2608.02474>Investigating Social Bias in Narrative Image Generationhttps://arxiv.org/abs/2608.01780>CopyCat: Improving Fine-Grained Subject Consistency in Subject-to-Image Models within Secondshttps://arxiv.org/abs/2608.00674>Diagnosing Under-Development of Irreversible Processes in Video Generationhttps://arxiv.org/abs/2608.00617>Where Does Generative Difficulty Reside? An Empirical Study of Target Representationshttps://arxiv.org/abs/2608.00626>UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reductionhttps://arxiv.org/abs/2608.01298>One-Sided Quantile Coupling for Flow Matchinghttps://arxiv.org/abs/2608.00978>MineGrad: Gradient Inversion Attacks on LoRA Fine-Tuninghttps://arxiv.org/abs/2608.01521>ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transporthttps://arxiv.org/abs/2608.00769>CoT-Edit: Let CoT Guide Instruction Video Editinghttps://arxiv.org/abs/2608.01113>Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detectionhttps://arxiv.org/abs/2608.00716>A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2 & Nano Banana2https://arxiv.org/abs/2608.01258>CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Modelshttps://arxiv.org/abs/2608.01644>DiffPrune: differentiable information throttling for token pruning in vision-language modelshttps://arxiv.org/abs/2608.01985
also ive learned if you dont have enough dialogue OR use the "0 to 3s: " stuff, you will get jibberish.but, it works: I need to add more actions or use more specific prompting.https://files.catbox.moe/4kirhb.mp4
>>109459200bruh I didun du nuthin :'(
>mfw MORE Research news>IDraw: Artist Verification from Digital Drawing Imageshttps://arxiv.org/abs/2608.01737>Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animationhttps://arxiv.org/abs/2608.01978>UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generationhttps://tanliming-daniel.github.io/UniMoCa>MoCRA: Mixture of Compositional Rank-1 Atoms for 4K All-in-One Video Restorationhttps://arxiv.org/abs/2608.01829>SPAE: Spectrally Guided Autoencoder for Pretrained Visual Latentshttps://arxiv.org/abs/2608.01306>Beyond Token-Level Cross-Entropy: Fréchet Distributional Post-Training for Autoregressive Image Generationhttps://arxiv.org/abs/2608.00562>Detail Continuation over a Trustworthy Coarse Scale for Autoregressive Super-Resolutionhttps://arxiv.org/abs/2608.01823>Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compressionhttps://arxiv.org/abs/2608.02109>SPARE: Structural Parameter-Free Affinity Regularization for Flow Matchinghttps://arxiv.org/abs/2608.01990>DAVET: Denoising-Aware Visual Evidence Trajectory Allocation for Diffusion Vision-Language Modelshttps://arxiv.org/abs/2608.01821>Decoupling semantics from vision: A framework for faithful visual-text compression evaluationhttps://arxiv.org/abs/2608.01848>SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generationhttps://arxiv.org/abs/2608.01977>Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learninghttps://arxiv.org/abs/2608.01314>ReACT-CLIP: Response-Aware Test-Time Defense for Vision--Language Modelshttps://arxiv.org/abs/2608.01067>RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocationhttps://arxiv.org/abs/2607.09757
>>109459245i look like this and say this
>>109459250KEKclassic projection it's some dumb brazilian posting a bunch of celeb shit on plebbit
>>109459243>>109459254>>109459262stop spamming nigbo
>>109459258>Assmongoy
>>109459243>>109459254>>109459262You seem lost
:::H3 SAMPLER SURVEY:::
https://litter.catbox.moe/49yq46pjc9dqywdx.mp4
>>109459294res_m / simple
https://files.catbox.moe/qyhjr5.mp4
Any of you guys tried referencing multiple voices from a single audio source yet?
>>109459151If the samples you use are clean it's pretty easy.I did some of that yesterday, I just could not get it not be blurry.
>>109459294dpmpp_2m / simple
>>109458949It's sandboxie, I have tried running it outside of the sanbox and now have possibly thousands of folders I will never remove, you used to be able to just plug in things and make it work, now it needs hundreds of nodes to generate a vagina. Fuck this sellout and fuck his niggerware.
>>109459319
minimax + starlight precise + flowframes =
>>109459327I don't get it
>>109459327qrd?
>>109459327So which custom nodes am I supposed to grab?
9 seconds, 0.9 MP, 2 references (1 audio, 1 image), 20 minutes in a 4060 Ti with 64 GB. With these times I think comfy needs a scheduler that can store prompts for doing all the inference at night ¿Does this exist?
anon your lora loader node?
>wan still in the OPcringe
>>109459327>So ComfyUICore is basically crippling first last frame modegive us context, which node is causing issues?
>>109459339>>109459327so if doing i2v or r2v, give very high resolution(about 4k res) reference picture?
>>109459371to the TE, yes
>>109459376But you can't even give the reference to the TE with the normal node. it's just "first_frame" with no control.That's why, What custom node lets you fix this?
which minimax version should i try with 24GB vram?
>>109459369he has custom nodes for feeding the qwenvl TE the start image as well, otherwise just as codex / claude to make it yourself
>>109459327>Garbage in, garbage out.Is this what pass for news? Fuck off.
>>109459391has nothing to do with quality / higher res starting image for the model, its using a high res version of the starting image for the TE
after some more testing I'm starting to realize anons here were righteven without the key loras for proper uncensored t2v, which I still suspect might be faster or more consistent, you could probably make do with just the ref model for almost any use case, it's really fucking good
>>109459319It does a great job putting the reference image character in the first scene, but then when it switches to the reference video in the second scene it just forgets them entirely. Will keep tinkering. Cool meek tho
>>109459001>>109459044These retards are already coping hard saying you don't need H3 loras lmao
>finally got my shit together and updated torch for the first time in year>nearly 60% speed bump with krea2jesus fuck
>>109459409LTX employee mad their model sucks, huh?
so RTX upscale helps to stabilize the generation artifacts a bit. very nice.
>>109459409by the time i get bored of H3 and need a lora a better model will have come outwe're accelerating
>>109459409bro just try it out. I was also a non believer but the reference workflow is really powerful. it's just a question of combining the right media inputs. yes that arguably takes more work but it's doable
I have a nice theory on why Minimax released such a powerful model, they're currently being sued by Disney, so they're giving a giant present to humanity so that we support them during the lawsuit, if the judge sees that people love Minimax and hate Disney it can move the needle
>>109459428the first part of your theory is a good one, but the simpler reason is now Disney gains nothing by suing them. anyone can violate copyright and deepfake common characters
>>109459200>noooooo you can't just legally make stuff I don't like!!!! ahhhhh the world is ending!!!!thank god I don't give a fuck about that place
>>109459428If anything this helps Disney's case... but now it can't be stopped.
What are your wait times currently with the lowest H3 shit you can do?With wan 2.2, I could generate 600x600 5 seconds in 30 sec, which was nice for prototyping. Debating whether trying H3 now or waiting until community optimizes it. I hear it's like 4 minutes for the shortest simplest video you could do
how much vram do i need to run that comfy workflow for minimax? i got 12 vram and 32 gigs of ramyes i know that i'm a vramlet but we all are unless someone here actually has a whole terabyte of it
I'll ask again, anyone genning h3 at higher resolutions and have the video change style and lighting completely?
>>109459445its worth tasting bro
Have we decided on the best stock wf for Krea2 and H3?I have a suspicion the comfy example is hot garbage for both
how can disney even sue minimax, a chinese company? they can just ignore them lol
I've 6 GB or VRAM
>>109459420>>109459426you can add as many references as you want you can't change the fact that the nsfw physics of the model is bad compared to wan and ltx loras. It's trash. Plus you can't combine animation style of smooth realistic and anime style. You have to choose between the two.
>>109459445performance seems to MASSIVELY vary across everyone's systems. Most of us probably haven't even been fucked to update torch and comfy kitchen in a year >>109459414and are probably using a fucked node setup that makes it slower (compacting the use sage attention flag WITH the kj node), and other fuckery. Your mileage will vary, but it's probably the most respectably performing video model we've ever had. LTX would always rape the shit out of my entire computer and freeze it or lag it, H3 doesn't even do that at a full 1.0MP 7s.
>>109459445and it sets your video card on fire too
>>1094594455 Sec is 50 seconds. 0.4mp, 5090, enough for shitposts/testing prompt.
>>109459443it really depends, I think Minimax is trying to prove a point, they release a powerful model, and if the internet doesn't explode in the next few months it meant that the fearmongering tactics were unwarranted, someone had to put the genie out of the bottle, see that nothing ever happens, and move on with our fucking life
>>109459467skill issue. boob and ass physics are top notch
>>109459474noice
>>109459470>LTX would always rape the shit out of my entire computer and freeze it or lag it,huh now that you mention it i'm able to do other stuff with my pc while waiting for a 10 minute H3 gen even though i am a vramlet, wouldnt ever be enjoyable with ltx
>>109459467I don't get dicks if that's what you mean. non issue
>>109459465know you're not alone in being a vramlet, fellow poorbro
>>109459413Why does this retard does not create an issue or directly dm/@ comfy or kijai in discord then, they are literally answering retarded questions and showing new stuff in discord like 12 hours a day.
okay this is a good one.the man Asmongold in image 1 is having a fight with the man in image 2 named Hasan, in the streets of New York during the day. Asmongold says "you dog shocking bitch!" to Hasan, and then Asmongold punches Hasan many times. After several hard punches by Asmongold to the face of Hasan, Asmongold knocks Hasan to the ground and Hasan says "okay stop it, I wont shock my dog!".https://files.catbox.moe/ndnj5r.mp4
>>109459498don't gen*
https://files.catbox.moe/4afl19.mp4
>>109459506what's with the low res on the faces? I thought H3 wouldn't allow you to generate low resolution stuff
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md>Example:> <Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.> <Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.Can someone explain to me why the Minimax documentation is referencing 4 pictures? The actual node only lets you select 3 ref pictures. Am I missing something?
>>109459531>Am I missing something?yeah, the 4th picture
>>109459531cant you just rewrite the node to add more references?
>>109459511damn that's impressive
>>109459531You add another reference it adds another slot
>>109459531Doesn't it add another empty slot when you fill them all?
>>109459511nice.
>>109459531>>109459540>>109459544Just connect more and it will show up automatically. Retards.
>>109459531Down syndrome levels of retardation.
>>109459557>>109459558Calm down you angry little autist
I hate how many retards new model releases bring.
>>109459565It's that and a blood thirsty gaggle of people who can't move on
this is insane I had no idea H3 could do porn>>109457874
>6 fingers
>>109459529doing 0.3mp to test/iterate fast
calm d-
>>109459565You're just autistic, anon. Lonely, autistic retards like you get angry at beginners because it's how your demented brain is wired.
>>109459565>noooo this general must be dead with only me and the voices in my head to shit on it!!!sorry schizo, but you lost
eceleb slop, amerimutts have truly rotten brains
Lots of insecure newfriends huh?
>>109459565kill yourself and it wont be an issue anymore
How good are video upscalers? I find that 0.8MP has enough detail, but I would like to blow some of these videos up.
>>109459531bwo...
>>109459629kek, bodied that freak
https://files.catbox.moe/7dy5cx.mp4Scuffed pearl with connies voice.
now we're talking
>>109459667big if true
>>109459593no one has a problem with beginners, everyone was once a newbie. we figured things out ourselves through experimenting and research.these newgens just want to be spoonfed, and it's a waste of time helping them. they'll make a couple of gens, get bored and jump to the next trend, whether it's a new video game or a marvel movie
>>109459667>4 stepsI always find that this is too low to keep the quality, maybe 6 would be the sweet spot
>>109459679Relax, autist. A beginner asking a beginner question is no reason to chimp out as much as you have been.
>>109459577you haven't been able to move on in 4 years now
>>109459679Why are we even discussing this?Why do we care about people being new?The only problem should be when they don't listen
>>109459608>eceleb slop, amerimutts have truly rotten brainsThe vast majority of hasans viewers are European. I guarantee you most of this eceleb shit is coming from a European or Paki.Now instead of crying perhaps gen some kino.
can minimax create interracial videos?>>109459696that's because Europe is caliphate land
>>109459679>these newgens just want to be spoonfedthey have the right to ask questions, and anons have the right to answer those questions or not, you having a meltie about it doesn't change anything
>>109459689Talking to yourself there bud?Trying to see if I can add details like cherry blossoms on the nails at higher resolutions
Anon please give the recipe to gen in 1mp, 15sec or more in minimax with a 4090 and 64 ram... I cannot gen without get out of memory... Why 5090 faggot can go for 15sec without problem?
>>109459727cute
>>109459700Yes it can do asian/white. No niggers, though they removed them entirely from the dataset.
>>109459731Skill issue. You do know that you can chain multiple shots together.
>>109459688>>109459691>>109459703a newbie asking a question is fine. but there are better places on the internet for those kinds of questions. these threads are meant more for professionals and advanced hobbyists. i'm leaning toward gatekeeping, because nothing makes a community more unbearable than the same basic newbie questions clogging up the threads. those belong on forums like reddit
>>109459667Lmk when ostris or any researcher or big finetuner tries to finetune the fuzziness and slow-mo crap away.
>>109459727the nails are pretty impressive
>a newbie asking a question is fine. but there are better places on the internet for those kinds of questions. these threads are meant more for professionals and advanced hobbyists. >i'm leaning toward gatekeeping, because nothing makes a community more unbearable than the same basic newbie questions clogging up the threads. those belong on forums like reddit
>>109459667Steps didnt matter with Minimax. I got same gen time with 20 / 15 / and 10 steps
>>109459756fuzziness is a sampling issue,slow-mo is a prompt issue.basically, git gud.
>>109459703honestly (heh get it), on /g/, if you're not using an LLM to solve/explain any issues you're havijng with software then you deserve to be poor
>>109459409This model is not anywhere near Sora 2 or Seedance tier, that's only Flux 3 which can do much more impressive stuff than this. I've been telling retards, but they don't listen.
I tested out Spectrum, the hands are a bit worse but apart that it's really clean, way better than EasyCachehttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3https://files.catbox.moe/4853id.mp4
>>109459753>gatekeepingThere are far more important things to gatekeep here and this general has been trying it's best to do so. More people the better this is a non issueAnyways 6mp or so takes 1 minute 20 seconds more detail is being shownTrying 8 next
>>109459789wish you had also put easycache in there.
So, this model is DOA after all right? Ryan is Minimax's CEO btw
>>109459789https://github.com/kijai/ComfyUI-SolAttn_tritonis worth it too
>>109459801>>109459245
>>109459787But, maybe if you're a VRAMlet you want to cope and say it's Seedance tier. Perhaps the weights can he saved, but as they are right now the model is still not even HappeHorse 1.0. The motion and picture quality is not fluid enough.
>>109459789the guy other guy lost his glasses.
>>109459797a bigger audience for your schizo faggotry?
>>109459798I don't need to, it's really a bad tool it destroys the audio
>>109459750Yes, but there are anons who claim they managed to do it using a 5090, and others using a 4070 Ti at 0.9mp and 15 sec. How? If mine won't let me do it with a graphics card that has slightly more RAM. I mean, I feel like I'm ignoring something. I'm using the ComfyUI workflow; often these workflows need improvements.
How is Minimax's music knowledge? Can I reference a popular song by text alone in the non-diegetic section and it will play something similar?
>>109459814it's unconsequential because I never asked for the other man to have glasses or not
>>109459789Post workflow plz
>>109459225I got a permaban from reddit for a musk dancing vid, just him eating that ice cream
>>109459822It's shit if you didn't see comfy's vid yet
>>109459829you just add the spectrum node
>>109459817people say this yet I find the audio to be fine. maybe it's because I'm using speakers.
>>109459822Local music was already saved by ACEStep. Just kind of sad we never got Qwen Music.
>>109459816I'm not here that often, You can always find a new hobby, care to show the class what you're working on?8mp results cherry blossoms are visible on the nails moving to 10 mp at 16 steps Does the text in my image upset you?
>>109459831>I got a permaban from reddit for a musk dancing vidwhy do you think I'm lurking here anon? 4chan is one of the last sites that allows a modicum of freedom of expression
>>109459831They did you a service honestly.
Internet Magical Girl!https://files.catbox.moe/wcjpls.mp4
>>109459855>moving to 10 mp at 16 stepsdon't even think you need to add more steps honestly.
https://huggingface.co/SexGod1979/NaughtyTimes_MiniMax-H3/tree/mainbefore the ceo finds it again
bajs?https://files.catbox.moe/lqevyy.mp4
>>109459878Style swings based on steps, sorry should have been clear. I have been using 16 steps from the start
>>109459876That's not AI right? You're just fooling around right?
>>109459801No is just a stupid faggotry to avoid government censorship, once the weight are open there are not a way to avoid it.
>>109459855now i'm really impressed by those nails, as a sdxlfren to this day i'm kinda jealous :>
>>109459801>nsfw tunewas it any good?
you guys are generating videos with AI, spending resources, while ppl are starving in africa????
given it's already pretty good I can't see it being that hard to train a nsfw lorahttps://files.catbox.moe/inhcur.mp4
>>109459906brb, batch generating Miku vids
>>109459853What was that other one that came out around the same time as the good acestep model? Was it HeartMula?
>>109459105Hi will h3 work on a 4070 (vram12gb), 16 gb ram, don't want to waste time downloading if it wont work thanks.
>>109459801Just don't put NSFW in the title of your lora, it's that simple.
>>109459879Is seem the audio becaume shit
>>109459908lora training for video models isn't that hard, only slow and expensive. a high-quality all-purpose nsfw lora should cost maybe $200? a project like sulphur or something could get it done easily.
>>109459908Its not THAT good for porn
>>109459917it's just that? like you can write "triple bukkake blowjob.lora" on civitai and it'll work as long as it doesn't have "NSFW"?
>>109459886messing with r2v, aika's short clip used as reference video
>>109459916yeah, you should be getting ~2 mins of get time for a 5sec video at 0.4mp
>>109459926not open weights yet
>>109459916works 4 me on a 3060 and 16gb ram
>>109459855This is one of the most impressive gens I have ever seen in my life. God bless you.
>>109459926who cares? dev is still not here and there's no way dev will be competitive with minimax
>>109459935>>109459937Thanks!
>>109459927Is there any open nsfw dataset or are we still at the "massive duplication of effort" stage of the "open weights but closed everything else" hellscape?
>>109459852My bisexual wife
>>109459926Bingo predictions:>SAFETYMAXXED>+30B params prunned>Nobody can run it>3x slower than H3
>>109459926Will the open weight version be as good as minimax place your bets now
is Illustrious over?
>>109459948there is no good open nsfw data, and honestly i don't even know how captioning video models works. it's not as easy as just throwing booru tags at the model. i think this is what sulphur is working on with their community curation project.
>>109459963not a chance, minimax is really uncensored, and knows every softcore/fetishes in the book (foot licking, armpit licking, actual french kissing...)
>>109459955if they actually give us the full weights with that sora 2 magic it has on api then I will be happy
it's 2026, nobody should caption their dataset with booru tags anymore. let it go.
>>109459985everyone should do both natural language and booru tags separately like how krea did it. It strengthens the effect of individual words
>>109459985those idiots will never give it up.>>109459994you mean anima?
>>109459970I still use Illustrious. I still can't let go of CLIP style blending.
>>109459970No
>>109459926BFL is known for cucking and lobotomizing their models so they produce body horror 80% of the time, just to make you cave in and buy their API services. Not excited for this at all
>>109459916How fast is your internet and what speed do you get from HF?
the reference model is wildthe man named Forsen in image 1 is standing with the man in image 2 named Hitler, at the Reichstag in 1940 Germany, in front of a huge crowd of nazi soldiers. Forsen says "Whats up, boys?" Hitler says "Hello Germans, I need you all to save Europe again. Especially you Forsen. Only you can save us Forsen." in a heavy German accent. The two men shake hands. the video is in black and white.https://files.catbox.moe/ismsfz.mp4
>>109459963>>109459926will do heckin wholesome astronaut dog on the moon dancing 10% better, zero nsfw, will be dead on arrival
>>1094599965' 9" woman
>>109460000krea lightly did as well, you just need to multiply the conditioning
>>109459801>>109459042bruh if civitai won't allow NSFW loras this model is in serious trouble, wtf are ou doing chinks, you're so close to perfection
>>109460020>heckin wholesome astronaut dog on the moon dancingWhy do you need more?
>>109459931Of course. Nobody would upload something NSFW without tagging it as such, so if it doesn't say NSFW, it has to be SFW.
>>109459794dumb bitch didn't even flip it
>>109460025but it's not deleted?
>>109459852Undressing fetishist here.New model is amazing at undressing/striptease. Still has some of the same old problems though, like shirts being split into two parts, gravity not always working properly (conditioning actually helps with the latter when you know what to prompt).You can now prompt for a character to put their hands behind them to unhook their bra before removal, and it actually works, complete with the unhooking sound.
>>109460025porn is illegal in china. It having nsfw in the dataset was a mistake
>>109458524>Do you guys have functioning ears? Can you not hear the glaring issues with the examples on this page? Sounds like dialogue coming through a tube or a toilet roll.I feed output files from DramaBox into CosyVoice for voice conversion to improve the audio quality.
>>109460020you missed the part where the dog has 5 legs and a flux chin
yeah so steps do have to increase the higher resolution you go with krea 2. The most I can stomach is 8mp at 16 steps when going over cfg 1
>>109460039>porn is illegal in china.yet Wan and Hunyuan have no problem letting westerners upload NSFW loras on civitai
my nuts are quakingwhen 4 step lora
>>109460052actually both of them had the first few taken down as well but then gave up
Ok I am a tremendous faggot and need an example of a successful prompt replacing a character in a reference video with a character in a reference image.
>>109460040How about you just use a better model like Fish S2 Pro?
>>109460012>buy their API servicesThere's an input filter. If anything, the API is even more cucked.
>>109460064let's hope it's also the case for Minimax, they're probably virtue signaling at the peak of their popularity and will calm down once people will have moved on to something else
>>109460052It's just a bit of PR management. The model just released, it would look bad in a headline if they didn't shut it down. give it a couple weeks and they won't give a shit anymore.
>>109460052The license of H3 forbids westerners from using the model
>>109459926Damn I wish I still had an x.com account so I could reply "Oh wow. Anyway..."
>>109460087Except if you're Canadian apparently.
>>109460087commercial use, individuals don't need a license to use Minimax
>>109460087Glad to be an Indian !
>>109460106Lets FUCKIN GOOOOOO!!!!!INDIA # 1 IN THE AI BIZ
>>109460087I think it's pretty funny that they actually had to make a statement about it. People are really dumb.
>547MiB / 24564MiB>run a gen>21387MiB / 24564MiB>Unload Models and Execution Cache>2347MiB / 24564MiBwhy is ComfyUI such a piece of shit at unloading models?
>>109460116Works for me?
The first 20 loras are almost always the runniest of shits you've ever seen.
>>109459926Keked again by a chinese lab. Definitely doing it on purpose.
It's actually 3 minutes at mp8 on my hardware but this is the cut off for small little detailsThis is absolutely amazing for a model on a single pass being able to juggle so fucking much. I highly recommend taking the speed hit and going over cfg 1. The chroma lora also does a lot of good work when properly adjusted with no actual quality loss.On to the next prompt
>>109460122First porn lora i've seen has excellent peepees and lady bits, tits are a bit bolt on at this stage.
>>109460138proof?
h3 is really something else when it comes to nailing voiceshttps://files.catbox.moe/h7ck21.mp4
>>109460147I don't think Jerry would say that.
>>109459643I tested seedvr video upscaler but it runs out of VRAM with 16GB at 1.2 scale with a 0.8 MP video. I'll try it in cloud when I get credits,In the meantime I am going to keep adding loop cuts to this until I get an hour of teasing.https://files.catbox.moe/iiev7g.mp4
>>109460146Follow some of the links posted.
>>109460147seinfeld probably not the best test for that tho, seems they made sure to include that extensively in the training
>>109460129That's good to know, thanks for sharing the results of your experimentation
>>109460147George's lips aligning to Jerry's dialogue at the start. Start again, cuck.
>>109460163>seems they made sure to include that extensively in the traininglol. It's just because the show has tons of episodes and they're always in the same places so the model can easily remember everything about it.
>>109460156tiled?
>>109460147Another Seinfeld h3 video gen? Jesus, why are you all so unoriginal, also, every guy who is posing seinfeld videos must be a 40 something guy
>>109460147I liked AI Seinfeld better when it was that low poly 3d version
>>109460207I like it better when George is naked.
dpmpp_2m_sde definitely seems like an upgrade over res_multihttps://files.catbox.moe/zn955z.mp4
>>109460193I tried to do a gen where the cast of Red Dwarf were at the table in the coffee shop and the 4 main cast of Seinfeld walked in shouting because their table was taken. It didn't recognise any Red Dwarf cast member apart from Kyrton, the cat literally was a black cat lying on the table
>Can't do celeb likeness>No gentials>Slow as fuck genDOA
>>109460222GO FUCK YOURSELF FAGGOT
>>109460147INT8?
>>109460147Are you the notorious 4chan reddit warned me about?
>>109460233you good?
I have 24GB VRAM and 32GB RAM, can I run H3? Are the half a dozen components of this loaded at the same time or in sequence?
>>109460260No. Don't bother. Its crap.
>>109460260yes
>>109460200The workflow has VAE decode tiled, I am going to try the trim video toggle and see if that works. I didn't read the tips of the workflow before.
>>109460260if you can run LTX you can run H3. In fact if you can run Wan you probably can run H3
You know what, I'm on board with ZITChad now. he's the ultimate retard filter.
>>109460260you might need to update comfyui and a bunch of random shit to get good speed but yes
>>109460260post workflow when you figure it out im cheering for you mate
>>109460233I think you forgot to take your meds today
>>109460260>>109460283Literal default workflows all work fine
>>109460222Specifically for 2D animation?
>>109459441it's illegal in america though (take it down act) and 4chan complys with american law
>>109460312TBD.I'm mostly looking at reducing artifacts. which for this specific gen, dpmpp_2m_sde is the clear winner.
>>109460318every 4chan thread eventually goes down, don't worry
>>109459159based jannies putting the schizo safety freaks in their place
What's the consensus regarding pruned vs not pruned INT8 diffusion model?
>>109459218>asmongoldlost all my respect for him after he decided to date a prostitute, it's Idubbbz all over again
>>109460356lossless.
anyone tested much with step count on minmix, 20 is comfy default but do you get a better video at 40, or do you need 40 if you go up to 6-10s videos
>>109459295you guys have to stop using litterbox, the links gets dead pretty quickly
>>109460361i don't believe this
>>109460368this. if it's sfw just cross post from /wsg/
>>109460375ok. have fun with your 33GB model.
convince me to update my comfyui setup to support h3im on amdi also have an arc setup for more suffering just in case
you guys use minmax workflow with upscale? I'm tired of generating low res after being spoiled by LTX
>>109460360pink_sparkles? I think she used to advertise her streams on here back in the day with hate baiting.
https://files.catbox.moe/cbmbkl.mp4
>>109460384??? LTX high res looked like this kek
https://files.catbox.moe/svyus6.mp4Kek this shit is retarded. Fuck this model
>>109460147>A jew insulting jewsnever happens, that's why AI is kino kek
>>109459913Yes, HeartMula. It wasn't that good compared to ACEStep 1.5
What are you guys using as a prompt to replace people in a video? The video always just ends up driving the still image of the person that I want to insert.
>>109460235yeah, out of the box, really>>109460245>>109460202https://files.catbox.moe/n16rg8.mp4
so it knows south park natively. t2v genhttps://files.catbox.moe/8u4g8d.mp4
>>109460431tell your LLM to prompt using the guides from the docs
>>109460446>>109460446
>>109460202>every guy who is posing seinfeld videos must be a 40 something guyI doubt that, Seinfeld is pretty popular in the zoomer community>t. zoom zoom
>>109460384 RTX upscaler at 1.5. Its acceptable
>>109460427ACEStep 1.5 XL got really good with LoRAs thanks to the Turbo/Base 0.3 merged model. With the right settings, better than Udio and Suno. On a class of its ownhttps://files.catbox.moe/j2l015.flachttps://files.catbox.moe/kwzpsz.flachttps://files.catbox.moe/7znul3.mp3https://files.catbox.moe/uv9k0b.flac
>>109460434you made me smile and for that i thank you. still have some issues with both characters trying to take the lines but ill forgive it.
>>109460407Were you born yesterday? Without a LoRA, H3 is way better than anything Wan or LTX would gen. Wan had genitals scrubbed out such that it would generate gore if you tried to prompt for them.
>>109460431<Video 1> is the source video for the editing task.Use the character from <Picture 1>, entirely replacing the original character in <Video 1>
>>109460485Look at what I just posted.Do you think that looks good? If you do, you may actually be retarded.
very light training. Needs undistill lora before proper training as it hurts quality too muchhttps://litter.catbox.moe/or1wyo8h856m9zjq.mp4https://litter.catbox.moe/a7krt97d89r296fy.mp4https://litter.catbox.moe/ao8ib8xk7m2t4hv4.mp4https://litter.catbox.moe/or1wyo8h856m9zjq.mp4
>>109460521https://litter.catbox.moe/222yihrgpmoqibrr.mp4
>>109460431There's a specific prompting format that is recommended. >https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md>https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.mdAsk your clanker to read these and translate your intent to a prompt in this format.
>>109460501>Look at what I just posted.Prompt issue, skill issue, you issue
>>109460485>replies to namefag trollWere (You) born yesterday?
>>109460534Or literally just do this instead and stop relying on clankers for literally everything >>109460499
>>109460499>>109460534replied here >>109460563
Is using a reference video supposed to make things like 10x slower than without?
>>109460599yeah
>>109460598just specify more in your promptreplace the [blonde/blue-haired/green goblin/whatever] in <Video 1> with the [dark haired sexy slime girl/osama bin laden/whatever] from <Image 1>
>>109460681crazy that we still don't have models that train inputs by an actual number.
>>109460695frontier models can do it
>>109460534I fed the instructions and 10 pages of my script to Claude, I have references for all the characters involved already. Here I come Hollywood.
>>109460710They use an LLM in the backend to rewrite the prompt. My 24GB is already maxed out. Might actually be time to get a spark.
>>109460521Use case for watching old men have sex with women?
>>109459926>and open weights coming soonthe fire rises
Is there any reason to go above nvfp4 for h3 on a 5090?
>>109463140why would you use fp4 over int8 convrot or fp8?