Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109465146https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
>>109466297thanks for the bake anon, godspeed
<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.the setting is New York City during the day. <Subject 1> plays a rock song on her guitar, then hits <Subject 2> in the head with the guitar several times, breaking the guitar. <Subject 1> steps into a cardboard box while looking shy.https://files.catbox.moe/3knoxq.mp4
>>109466320valid crashout desu, he tried to swallow her with his body or something!
>>109466297>/g/ still doesn't have sound in the year 2026come on this is unnaceptable :(
>>109466332Literally 0 reason not to have it sitewide
>>109466334they could prevent 'muh jumpscares' or whatever other retarded motivation they have by having the videos start muted by defaultbut it's /g/ we're talking about :)
>>109466338>by having the videos start muted by defaultWhich is already how it works on the boards that have it. Actually less than zero reason because it's extra shitcode that makes the site worse for no fucking reason whatsoever.
I'm goin for 14 seconds 1MP, might be a while
>>109466334the new site janny broke the code for timestamps and tripcodes. i don't think he knows how to enable permissions for sound
Another thing to work on: environment consistency over longer times / multiple shots.
im surprised how clear the audio is in this model. and this was at 0.3 and still pretty clear!<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.the setting is New York City during the day. <Subject 1> plays a rock song on her guitar while standing near <Subject 2>, then hits <Subject 2> in the head with the guitar several times. <Subject 2> falls on the floor and says "bocchi, I cant breathe!"https://files.catbox.moe/2ugwwz.mp4
https://github.com/lihaoyun6/ComfyUI-MiniMaxH3-Cachesuggests it's better than easycache. no tests/doesn't elaborate.
>>109466363>no tests/doesn't elaborate.I hate when that happens, or else it's some vague posting king, or else it's some giant LLM wall of text that has nothing of substance in it
>>109466357Kino
>>109466362also, 185 seconds with 0.3mp/10s and the patch sage kj + spectrum nodes, this is my setup in the reference workflow:
>>109466352>i don't think he knows how to enable permissions for soundthere's no excuse we live in the LLM era he can ask claude Idk
>>109466375replace the patch sage note with MiniMax H3 Mem Eff Sage Attention Patch node.
with this reference model you dont even need loras although obviously loras are best for 1:1 characters. but look at bocchi go!https://files.catbox.moe/2bznao.mp4
this model is surprising me. i can actually prompt for the tank commander to peek down into a detailed tank turret with crew members doing their job
>>109466391ah, I have that in my i2v model workflow but forgot to add it in the reference one, ty anon
>>109466393>with this reference model you dont even need lorasIkr, it's amazing at keeping the character consistency, Klein dreams of being this good, people on twitter have to bully the Minimax CEO into making an image/edit model, he would nail that shit too
>>109466348I hope it's just the preview fucking up but it's showing nothing but noise
>>109466401adding that to the reference workflow speeded things up as well, I guess it optimizes the vram or memory or whatever. pretty cool how rapidly people are optimizing this, then again not surprising since it's a genuinely good model and not slop.
I havent gen'd since an old wai SDXL a few years ago and was wanting to pick up some current models. Is krea2 for realistic and anima for 2D the current meta or ?and do I really have to make an account to download krea2? I cant find it on civitarchive, and its cucking me on HF
https://files.catbox.moe/zonwbf.mp4punching aside, I love how this model can actually make realistic sounds, the guitar sounds like a guitar etc.
LTX and wan might as well just close up shop, and BFL, other than klein edit 9b and krea 2, wtf do you need other than thishttps://files.catbox.moe/b5kvq7.mp4
>>109466393Can you give multiple references for the same character? Like a front view and a back view
https://huggingface.co/Kijai/MiniMax-H3-TAEThis shows noise for me. kjnodes is up to date. anyone else?
>>109466430yes, front back and side keeps it from guessing
hahaha, this is a whole new level of shitposting potential.<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.the setting is New York City during the day. <Subject 1> plays a rock song on her guitar while <Subject 2> plays the drums, <Subject 1> is singing "he can't breathe, he just wants some fent, who knows where it went, George Floyd.".https://files.catbox.moe/xr8sv0.mp4>>109466430yep, you can give camera directions and even specify shot by shot, ie 0 to 3s: view from behind subject 1, etc
>>109466437The structure is even more sophisticated in the example at the bottom of the guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
bocchi concert part 2, I didnt even prompt the camera directions.https://files.catbox.moe/qtdmje.mp4
Okay, minimax is pretty fucking greathttps://files.catbox.moe/ysfmy6.mp4
Just deleted wan and ltx lol
>>109466408Okay it came back, kind of it's still noisy but there's something in there. I don't think the preview was designed to deal with latents this big
>>109466486based
>>109466313I dunno how it performs but you can combine sol and sage.They modify different things and are in theory compatible.
getting some good music from this
Continuity is difficult. I think I will need to try feeding it a whole storyboard.
Any resources to getting started with minimax h3? I never opened comfyui in my life.
>>109466560Click the template button. Load the minimax h3 t2v template. Download the models. Lurk these threads for tips
<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.this is the key to getting interactions in the reference workflow. then just say <Subject 1> or 2 does (whatever).https://files.catbox.moe/z0543u.mp4
>>109466558we don't have the SCAIL2 type equivalent to this yet but even just passing references to past scenes and characters and making use of 20s+ segments should do a lot to get "continuity"?
>>109466572Does it still work if you do something like "<Subject 1> is the anime girl in <Picture 1> and <Picture 2> and <Picture 3>"?
>>109466560as the other anon saidif you prefer a simpler UI for now wan2gp also supports it, but you won't be able to use various of the nodes people here use.
KINO<Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.the setting is the city Tattooine in Star Wars. <Subject 1> walks towards <Subject 2> and says "you tried to sell my lightsaber for fent, and now you must die.". <Subject 2> holds up a lightsaber with a red beam and says "I WILL get high, Skywalker." in a black man's accent.https://files.catbox.moe/9wmqx4.mp4
>>109466588once you establish picture 1 (image 1) as subject 1, always refer to them as subject 1. like a variable or whatever. the model is assigning <subject 1> to that image. if you use multiples for one name it might get confused
great more floyd gensreally funny lol!!!!!!! so much creativity!!!!!!!!!!!!!!!
>>109466596I hope you're making other stuff too, you don't have to try to impress us
>>109466600I meant for using multiple reference images for the same character. I've seen the version where you specify that the face is from <Picture 1> and the body from <Picture 2>, but I was wondering if there was a way to get it to use multiple images for the entire character.
>>109466596https://files.catbox.moe/v2qwhg.mp4
>>109466608In this house George Floyd is a hero end of story.
>>109466608we tell stories of our heros todays so that they may be remembered as gods tomorrow
>>109466613I am, Miku and Floyd are my "hello world" test case just to test concepts.
>>109466618Better than he ever was
>>109466391nta but can't both not be used? hmm I'm currently using both do they conflict or something? I'm always looking to squeeze a bit more out of my shit hardware.
>>109466415ok well back to wai SDXL i guess. anyone got a link to the local tag DB with autocomplete?
>the sand people will be back and in greater numbers
>>109466638XD
>>109466638go back dumbass
i still dont know what wai stands fort. newCHAD
>>109466642>>109466644:( not very nice!
>>109466655Waifu
shieeeeeeeeeethttps://files.catbox.moe/4c7nwn.mp4
>>109466558>>109466574Chaining shots together is an option. I still need to git gud at shot cuts and transitions.https://streamable.com/l6j9uy
>>109466655Lora shitmix of obsolete model popular with Bharatis and Hispanics. Nothing you really need to know newGOD
>>109466636You just need one and that one is optimized for minimax.
>>109466678kk
Got minimax h3 and it seems time to get back into this again.Still hate writing prompts.
>>109466678Well whats the goto for animoo then? anima?
>>109466663We'd have killed back in high school for this tech
>>109466703I am rawdogging short lazy prompts and it is working good compared to how low effort I am with it.Just Video:1-2 sentence description followed by Audio: 1 sentence description, no scene/second breakdowns.
>>109466711Yeah.
i've been enjoying H3 so much that i downloaded it a second time
>>109466728don't be greedy. leave some for the needy
What do you think about integrating RTX Super Resolution in your minimax workflows? The video looks good to me, but I don't know why it failed to reintegrate the audio.
>>109466727does it need a lora for the sexo stuffs?
>>109466664if you are serious about slopping(juding by your shit i assume you are) you might want to look into divinci resolve and more traditional ways of putting together longer videos. probably speed up your work quite a bit and give you more creative space to work in.
>>109466744It knows most sexual concepts, it just struggles with uncensored female genitals.A lot of body horror if you try to generate vaginas
>>109466664>>109466574I'm generating in single 30 second chunks at a time, but it already forgets the structure of a complicated environment when it switches scenes within the same generation. I think it might be better to do only single cuts per generation and splice, but that would be a lot more work.
>>109466760It wasn't trained with 30s cuts, I don't know what you expected
>>109466744I generated a lot of hardcore shit with it without loras, so no.
Black Forest Labs insider here. We're fucked.
Anyone get this to work on AMD?
there is a latent upscaler, it has some problems thoughhttps://github.com/Tr1dae/ComfyUI-MiniMaxH3_LatentUpscaleri seems with this model earlier steps do more with the video and latter ones with the audio.
>>109466772on god. isn't that shit releasing in a few days?
>>109466767Pretty but needs a "TRVTHNVKE" somewhere, probably on the wings.
>>109466769from what we know, the total video length is not a factor, but the length of an individual cut within a sequence may be limited.
>>109466574can't you just like do 30 seconds, then feed it the last 10 seconds of the first video and extend it 20 seconds and repeat? I thought the model was capable of that, i'm guessing you would also need to feed in character reference sheets and props. I haven't even begun to test those features out yet.
Jesus fucking christ, i cannot get enough of this modelhttps://litter.catbox.moe/p05yi0i570j87o02.mp4
>>109466767Shows how it struggles with the sequence of events, with the ground already burning in advance.
>summary: mayli from facialabuse.com is sitting in a cafe alone. watching her own video on a laptop.>the camera start with medium shot on her sitting on the table. after 2 seconds it shows her watching video on laptop, in camera third person watching her.h3 t2v doesn't know who is mayli. trash model
>>109466792Yes, but the problem is that even within the first 30 seconds it already loses track of where it was. Doing one scene (cut) at a time and editing together seems like the safe way to go.
>>109466767i haven't tried any nuclear kinos yet with it
>>109466760try timestamps like 0 to 3s: (action) for every shot up to 30s, it may work but idk, works for 5-10s
>>109466560Comfy is really not that hard to operate, just use template workflows and remember the general control flow, then play with nodes and see what happens when you make a mess out of them.
>>109466811anon, there is an entirely separate reference model...just plug them in and say <subject 1> is the woman in <picture 1> and you are good to go
this is wild, the possibilities are endless.multi reference drifting!image 1 is anakin, image 2 is floyd, image 3 is a random background from jojo part 7.The video takes place entirely within the setting, and environment shown in <Picture 3>.<Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.<Subject 1> is facing <Subject 2> who is 10 feet away and says "you tried to sell my lightsaber for fent, Floyd.". <Subject 1> takes out a green lightsaber and <Subject 2> takes out a red lightsaber, and <Subject 1> and <Subject 2> have a high intensity lightsaber fight.https://files.catbox.moe/xtmxgc.mp4
>>109466834reference background image:
>>109466844cute!
>>109466827>>109466569>>109466593Thanks for the tips, I had no idea comfy had templates. I'll download it right now! I'm excited to test it
>>109466486WAN yesBut for static animations LTX still good
fent ball run. I added a music track from the show.The video takes place entirely within the setting, and environment shown in <Picture 3>.<Subject 1> is the man in <Picture 1> and <Subject 2> is the man in <Picture 2>.<Subject 1> and <Subject 2> are riding horses side by side. <Subject 1> and <Subject 2> ride their horses very fast as a crowd watching in the stands cheers, and <Subject 1> crosses the finish line first, with <Subject 2> close behind him.https://files.catbox.moe/9i347q.mp4
>>109466767Some anon's mom really shut it down geg
2026 truly blessed, we got krea and minimax and it's only august
>>109467036we also got anima in january iircand we got the reject to mostly stop trying to shill his trash
Have any of you cucks mastered ref prompting in Minimax yet?
It feels like I was meant for more in life
>>109466943How are you getting such a clean video at that res anon?
>>109466420>other than klein edit 9b and krea 2, wtf do you need other than thisAnima
>>109466297Does anybody know the workflow or steps needed to get Minimax H3 to gen as crisp as this vid? Ive been trying things all day and no dicehttps://files.catbox.moe/pj2jjl.mp4
>>109467087close up being crisp is expectedfor non-close up the only thing i can think of is higher sampling resolutionfancy postprocessing video upscalers might help as well (not sure)
>>109467077It's the default workflow with sage attention.
>>109464587Thanks i ended up doing this>tfw faces visible for the first time :o
Is Qwen3 obsolete? Should I get rid of the models?
>>109467098Yeah Im trying upscalers atm, but ive been hitting dead ends for 7 hours
>>109467119Yeah, especially qwen3vl_32b
>>109467060Yeah and your mom is incredible; Rimming niggers, Asian gang-bang, pool sized bukaké she can't get enough.
Any fag wants to test res_2s beta57 version of any of his already posted clips for comparisons?
Btw - you can somehow fit int8_convrot qwen text encoder in 32 + 64, but with UnloadModel node
reference model is so fun.The scene is set inside the wooden horse-cart from <Picture 1>. The original blonde Nordic character from <Picture 1> must remain seated on the bench exactly as shown. Introduce the new character from <Picture 2> as <Subject 1>. The blonde Nordic character from <Picture 1> is <Subject 2>. Position <Subject 1> sitting directly beside <Subject 2> on the same wooden carriage bench. <Subject 1> turns to <Subject 2> and says "where is the fent at, nordic brother?". <Subject 2> says "we only have skooma" in a swedish accent. The carriage bumps and sways down the dirt path, causing both characters to physically react to the bumpy ride together while looking at each other.https://files.catbox.moe/5wfgc1.mp4
how clean do voice samples need to be to be usable?
Can confirm MiniMax H3 Sigma Shift is a no-op if you apply it to guidance only, it uses the default values in that case.An anon asked a few threads ago, just confirmed my guess now
>>109467153Not that clean. Im making a zero shot cloning workflow atm, its really good so far Ive tested
>>109466375Try it with a video ref and watch it drag
>>109466608That's SAINT G. Floyd to you
>>109467128I use "4x_NMKD-Siax_200k.pth"You should try it and show us the result.
>>109467185more reason to hate gingers
anyone know, what existing TV media does minimax have really good knowledge about apart from The Office and Seinfeld?
>>109467153not at all, I've used clips with background music and it turned out pretty good. Granted it wasn't a voice I've heard a lot and could recognise as perfect but it was well in the ballpark
>>109467206It does better if the voice is linked in the prompt to the character that you're generating also
Why is video generation so complicated? I just installed comfyui to generate porn on my new GPU but wtf how many tools, plugins, extensions, models, hours of learning do you need to get something decent?
So when will minimax be able to do nsfw? Will the creators successfully shut it down?
>>109467185>didnt bite the connectorsmart cat
>>109467212it works without doing that?news to me
>>109467218If youre using an audio reference, why wouldnt it?
>>109467225because you haven't told it to?I would suspect it works better if you do
>>109467213Beg coding with ChatGPT doesnt help that much, and people are coy with the best workflows sadly. Fair enough, but very annoying when people send you down bad paths for shits
Do people use "director" nodes now or are they a meme?
>>109467235some kind of boiler plate for the syntax?
>>109467238kek
>>109467213Lot of things go into making even just porn. You also need to learn about cameras, lenses, lighting, color, perspectives and so on. To make something good you still need a film school crash course basically.
just watch porn and learn the angles
How do I prompt H3 for POV video???
>>109467252have you tried POV
<Subject 1> is the man in <Picture 1>. The setting is an in game perspective of icecrown citadel in the game World of Warcraft. The player character is a warrior with a large sword. The warrior charges after a raid boss with the appearance of <Subject 1> and hits him with the sword several times.im dying. and that's actually part of the raid. gonna try higher res next.https://files.catbox.moe/cjctnw.mp4
We can finally get the collabs we were robbed ofhttps://files.catbox.moe/ckfdr2.mp4
virgin porn watchers vs gigachad sex investigators
Whats the best workflow atm?
>>109466793What resolution is that video?What is your hardware?How long did it take to gen?I2V or R2V?
>>109467255yes
>>109467265currently this node setup seems to work better than easycache for me at least:
>>109467282Cheers
>>109467256KINO, now this is classic wow.https://files.catbox.moe/jjk9of.mp4
>>109466391I see just ltx and wan mem Eff Sage Attention Patch node.updated comfy yesterday
>>109467256I'm noticing those machine elf hexagons all over my gens as well
>>109467294update kj nodes
>>109466415I use Krea 2 for anime too. Go find the krea 2 fp8 scaled version on HF or get the int8 convrot but I couldn't get that running.
https://i.4cdn.org/wsg/1785931146429739.mp4>>109467242yes, they look like audio/video editing timeline UI.
>>109467238Reminds me of a web game I played as a kid called Fly Like a Bird 2 where you shit on people as a bird in a city. Pretty insane that it ran in my browser on my 2007 Pentium >>109467282Nta but thanks. Will try spectrum over easycache shortly. Minimax H3 is so good I actually want to make sfw videos with it to see what it can do. It mogs seedance 2.0
>>109466338Could literally just have a sound icon on the thumbnail for the "jumpscare" warning
360 seconds for 13 seconds video at 0.6 mega pixels1 hour = 3600 seconds3600 / 360 = 1010 gens every hour10 x 18 = 180180 gens / day What can i do for 180 gens ??We need more optimizations. I wonder in time Minimax will be optimized so you can gen under 250 seconds ??
lmao, not exactly what I wanted but it DOES know a lot of games.<Subject 1> is the man in <Picture 1>. The setting is an in game perspective of the game super mario 64. Mario is running through a level with grass and coins. Mario runs and leaps on <Subject 1> and a coin appears after Mario jumps on his head.https://files.catbox.moe/rrkqla.mp4
>>1094672670.5 mp3060 12gbaround 13 minutesi2v
Can you run this on 8gb vram 32gb ram on Windows?
everyday postinghttps://files.catbox.moe/538de0.mp4
>>109467314>Fly Like a Bird 2i also thought that exact thing, but i didn't say anything since i didn't expect anyone here to know about it
so long gay bowser...https://files.catbox.moe/hokybc.mp4
>>109467326thanks, anon. That's about what I'm getting, as well
>>109467315>Could literally just have a sound icon on the thumbnail for the "jumpscare" warningAny of the top 5 AIs right now would be able to fix all of the sites problems and add all the missing features in like 4 hoursAt this point 4chan iws defined by the fact that you can only post 1 image per reply
>>109467192Kill yourself
>>109467235It's cool what you can achieve with noodles, but If you are going to have to edit something from separated clips you should use a proper editing software, there are free ones like Davinci Resolve, I personally use Shotcut. There might be niche situations where a director workflow can achieve what simpler workflows but outside of that simple is better.
>>109467185Always an orange fucker
>>109467352gm saar
loli
>>109467392I think the director nodes are good for filling in the gaps. Like using the end frame of clip1 and the beginning frame of clip3 to generate clip2 with FL2V.
Been traveling the last few days. So is minimax gods gift to coomers or another nothingburger?
rate my songlate 80s metal>>>/wsg/6207803
>>109467433I haven't genned a single coom clip. It's fun model though.
>>109466297Help me connect Spectrum nodes because i have no idea what i am doing here
it knows EVERYTHING.<Subject 1> is the man in <Picture 1>. The setting is an in game perspective of the nintendo 64 game the legend of zelda ocarina of time. Link opens a door in a dungeon and on the other side is <Subject 1> who is 10 times as tall as Link. Link slashes him several times with his sword and then <Subject 1> disappears, and a heart container falls on the ground where he was standing.https://files.catbox.moe/9wjqug.mp4
>>109467258why won't they die?
>>109467435i rate it k for kino
>>109467433>>109467436https://files.catbox.moe/gcupql.movMate made this. Dont know what the work flow is, but it seems like a gem
>>109467444
>>109467474Thanks bro
>>109467408good morning saarshttps://files.catbox.moe/lcvwdh.mp4
>>109467472Oh it doesn't autoplay .mov files
>>109467472Based and goonerpilled
>>109467502https://files.catbox.moe/fsytg4.mp4Here we go, Misato titties
>>109467472I've seen plenty of good goon gens made with it.
the reference model can do style transfer too. this is KCD for example.
>>109467507What I havent seen is the workflow that gets it to this level
>>109467379Can I fuck your mother first? And your sister? I already did you father.
>>109467483is spectrum actually worth it? Please report back, anon
>>109467517You are brown
like this guise?H3 mem eff sage attention before spectrum?
>>109466953>>109466827>>109466569>>109466593https://files.catbox.moe/bac7ag.mp4Ha, it worked.
>>109467515is that musa?
>>109467525Yup, that's why you father has aids now. He was kind of a slut you know?
>>109467534Now unpack that subgraph and start adding optimizations to generate things faster
The default shift/scheduler/step count combo feels ass to me.What are (You) using?
>>109467546>unpack that subgraphHoly shit, that's a thing?! What the fuck. The rabbit hole goes deeper than I thought.
>>109467534congratulations newGODpost some kino in the thread
Henry and Floyd's adventure:https://files.catbox.moe/b1au0e.mp4
how do I load loras for minimax? which node?
i been tweaking the same prompt for 5 hours now and i think it regressed
>>109467572>xe overfit on local prompt minimangmi
>>1094675722 hours but still having fun.
>>109467572Im at 7. As soon as I find the right one, im spamming it in every thread
>>109467572you need to use an llm for optimal h3 prompts
>>109467570
>write elaborite prompt using the guide>model completely ignores the prompt and does random bullshit>?????>Wtf? is this normal?
I used pizza tower as a style reference and this happened:https://files.catbox.moe/h5mkda.mp4
>>109467582the llm isn't good enough for kinosovl
>>109467600maybe it needs better director.
>>109467590Don't waste time writing it yourself. Get a clanker to do it.
Are we entering a local golden age?? Seems like APIs are getting raped in the ass, with OpenAI having to drop prices massively in response to chinese models, and xi has openly endorsed the open weight approach. qwen is also releasing their Max LLM locally, something they haven't done before.But where is Qwen Image and Wan???
>>109466999trips of double ice coffee
>>109467620>Seems like APIs are getting raped in the assWhen deepseek v4 pro drops we gonna see anal gapes so big that no goatse could come close.
>>109467620looks interesting
>>109467585won't let me hook italso the h3 mem eff sage patch is not working
>>109467521Just tried it. Nope. Probably my workflow fault. But i heard Spectrum lowers quality
>>109466811try kelly balthazar
Fuck off back to your containment thread retard
>24GB vram>64GB ram>R2V>736x416>177 seconds to genaudio: https://files.catbox.moe/9b5g8u.mp4
>>109466703Just feed the H3 prompt guide to your clanker and talk to it how you want your scene to play out.
>>109467527order doesnt really matter lmao
Reference model is good as a whole but wtf is up with using video references? I haven't got it to work once so far even with your basic jeet character replacement scenes.
>>109467527is the spectrum apply node snake oil or legit? i don't trust anything that isnt from KJ
>>109467668If you're having prompt adherence issues you need to adopt the official format.https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.mdI recommend plugging the guide's source into your favorite AI assistant (e.g. Fable) and asking it how to properly prompt for your desired result.
>>109467672absolutew snake oil. even easycache is better. you just should wait for KJ to wake up.
>>109467435impressivebut too noisy and more like early 90s
>>109467679I have already tried doing so, and I'm either getting weird outputs with my <picture 1> subject or getting <video1> thrown right back out at me. I've tried using it for both NSFW and SFW clips and can't get anything out of it.
fuck im a total promptlett barely able to gen images and trying h3, the audio its putting out is so fucking cursed kek. this is alot of fun though, 4:3 at 0.2mp genning 10s in under 2min just to practice prompts
https://github.com/lihaoyun6/ComfyUI-MiniMaxH3-Cachethis does not effect audio
>>109467691I do feel it's quite finicky so it may not be just you.
style transfer test, game is resident evil 1 original: pretty cool that it works.<Subject 2> is the man in <Picture 2>REFERENCES:Use <Picture 1> exclusively as a visual art style, texture, and color palette reference. Use <Picture 2> exclusively for the facial identity and clothing structure of <Subject 2>.INITIAL PLACEMENT & STYLE:The video begins with <Subject 2> positioned standing on the left in the retro video game scene shown in <Picture 1>. <Subject 2> is fully integrated into the scene from the very first frame, completely transformed into the exact pixelated, retro art style of <Picture 1>.STYLE RENDERING:The entire video must be rendered completely in the exact retro video game art style shown in <Picture 1>. Apply the pixelated textures, low-resolution aesthetic, specific color grading, and retro rendering artifacts from <Picture 1> globally to every asset in the video.SHOT & MOTION: <Subject 2> is walking around a large mansion, towards the door at the left of the room. <Subject 2> is fully transformed into the retro video game art style, looking like a playable game character.using a template from google ai search, seems to work.https://files.catbox.moe/rxpjl9.mp4
https://www.reddit.com/r/StableDiffusion/comments/1veb4bn/i_created_a_sysprompt_for_minimax_h3_to_emulate/https://gist.github.com/Naxdy/43b7422a1e4a79fb8b0489c6c39eaaceFor the local bros this sysprompt.md works perfectly with Gemma 31b q4 to create H3 prompts from just single line prompts
What are the good prompt settings for animating figurines in minmax? Should it be 3D CG or Live Action? Any other parameters?I tried live-action and it looked a little weird and the output sometimes had errors.With wan2.2 it just werked pretty well out of the box (webm related).
>>109467642>also the h3 mem eff sage patch is not working it needs minimum i think 2.0 Sage Attn. That's my problem with it and I cannot get 2.2 working on Linux, I may have to compile it but im not in the mood for fixing shit if it goes wrong today. I don't even have GCC installed (I do but apparently it doesn't exist ...)I got the 2.2 version from comfys repository, that didn't work for me.
>>109467709Oh so I need sage 2.0tried installing it but failed
>>109467721just give them an action and specifically say to keep them in the style of an anime figurine or sculpted model or something.
>>109467721Throw the prompt guide into a clanker and tell it to rewrite your prompt in that style, then just tell it what you want.https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.mdhttps://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
same re1 style transfer prompt but 10s test:https://files.catbox.moe/do0ecs.mp4
>>109467634>bruh is so crazy you see video of sinde sweeny kung fu fight batman! and then everyone forgets the model exists in a week because normies aren't wasting 20-40 dollars to gen a few seconds of video.
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731https://huggingface.co/moonshotai/Kimi-K3
>>109466297Can Minimax H3 run on a 4060 8GB? Whats the hardware requirments?
>>109467750same prompt again but with link to the past:https://files.catbox.moe/fgmy5f.mp4
>>109467747Man I don't know about using an LM to write my prompts, I've always written them by hand because I desire control.Like, I know what I want, but the guide says you should specify whether the video is 3d cg, live action, 2d animation, etc. Moving figurines is a bit of a tricky one to quantify.
I'm trying both qwen3tts and fish s2. I think claude has cucked me and I can't get what I'm after.I want to take a voice, preset and/or voice cloned and customize it's emotions for sentences. Do I have to git gud or is there something better?
genuinely confusing how we got something like H3 released open weight. this model has zero limits
I used the word fentanyl in the prompt and the video became worse quality, sort of wobbly. I stopped using it and it became better.>>>/wsg/6208528>>>/wsg/6208526
>>109467784alright that one made me laugh
>>109467828yeah, it's even better in cunnygen than Wan 2.2. i'm kinda blown out
we have AI salty bet now. just pick 2 images and have them fight. make the prompt ambiguous so the AI decides the winner. "one of them falls down", etc.<Subject 1> is the anime girl in <Picture 1> and <Subject 2> is the man in <Picture 2>.the setting is New York City during the day. <Subject 1> is having a fist fight with <Subject 2>. <Subject 1> punches <Subject 2> very hard several times, and <Subject 2> falls to the ground and says "I can't breathe!". <Subject 1> uses a peace sign gesture with one hand and smiles.https://files.catbox.moe/c90sc5.mp4
I NEED BETTER MINIMAX WORKFLOW>I NEED BETTER MINIMAX WORKFLOWI NEED BETTER MINIMAX WORKFLOW>I NEED BETTER MINIMAX WORKFLOWI NEED BETTER MINIMAX WORKFLOW>I NEED BETTER MINIMAX WORKFLOW
>>109467862use dasiwa's workflow from civitai
its SOOO fucking fast nowhttps://github.com/lihaoyun6/ComfyUI-MiniMaxH3-Cachehttps://litter.catbox.moe/p6c7n2nu49gwhi9g.mp4https://litter.catbox.moe/bsrgokundp30mjc0.mp4https://litter.catbox.moe/zjqkhr6vvsp27fmn.json
>>109467866100%|| 20/20 [02:03<00:00, 6.18s/it][MiniMaxH3-Cache] Skipped 10/20 steps (2.00x speedup).That was 2MP btw
>>109467862just wait until we have nsfw and turbo loras
https://files.catbox.moe/svh1zf.mp4
PSAhttps://files.catbox.moe/9pa6ao.mp4
asuna from blue archive fights crime:https://files.catbox.moe/rax8j0.mp4
>>109467866>>109467871Tutorials ??
added this lora loader, hope i added it right
>>109467898literally has a WF and link to the new node
Convrot is pretty much a necessity right? I tried the non-convrot models and motion definitely felt stiffer, but that was back in only my first handful of prompts in Minimax.Do all of you guys run convrot models?
>>109467866>2 days ago>zero activitysnake oil, kills the quality
>>109467911bf16
>>109467911no reason not to use cr
>>109467911yeah even with klein edit and krea it's faster and bf16'ish quality, fast and quality so why not
>>109467916it truly does nothttps://discord.com/channels/1076117621407223829/1534342635597070466
>>109467916it's a chinese virus, don't download it
>>109467928
>>109467925>no reason not to use crWell, there is. The non-convrot nvfp4 model is like 8gb smaller.
>>109467936its literally the bandoco serverhttps://discord.gg/4gVXrgsxWits the "speed up H3" under H3 resources channel
>>109467794The datasets behind the models are all tagged by LLMs, so its not a bad idea to use one to fluff up your prompt
>>109467960KEK
I just deleted all my wan models and loras.They served me very well, but it was finally time to let them go.
>>109467966how do i unsubscribe from this blog
>>109467954Ok, and what do you recommend for an abliterated/uncensored prompt writer? I generate lewd and I am not tolerating rejections.
>>109467978gemini 3.5 flash
>>109467943and the q1 is 3x+ smaller than that, almost like quality below int8 goes to absolute unusuable shit that no speed increase matters
>>109467960average arch user
>>109467982>cloudno
>>109467892
>>109467983then why comfy defaults workflow to nvfp4 qwen?
lmao you can literally plug in any image and it will work, kek<Subject 1> is the dog in <Picture 1> and <Subject 2> is the man in <Picture 2>.the setting is New York City during the day. <Subject 1> is having a fist fight with <Subject 2>. <Subject 1> punches <Subject 2> very hard several times, and <Subject 2> falls to the ground and says "Okay I wont shock dogs any more!".https://files.catbox.moe/e927ds.mp4
>>109467983The question was not related to quants, it was related to convrot.minimax_h3_fl2va_pruned_nvfp4.safetensors 12.5 GBminimax_h3_fl2va_pruned_nvfp4_convrot_int8.safetensors 20.1 GBSame compression, one has convrot and is 8gb bigger.I'm just saying there technically is a reason to use the non-convrot, but I'm curious what kind of effect it has.
>>109467978Gemma 4 heretic or uncensored models I guess
so is stable audio better than acestep ???
can we have separate vid gen and image gen generals?
>>109468027MoE?
>>109468066>>109468066
>>109467978DS4 Flash. Just write an appropriate system prompt, together with the mass of text from the prompt guide that will suffice to make it comply even without a prefill. If it somehow still refuses whatever abominable prompt you give it, just add a prefill. Without closing the think tag so it can still continue thinking.
>>109468002for h3 the text encoder is huge and 4 bit quants are ok while being small which most people need since they are v/ramlets, in basically every other case ever those are dogshit quants that should never be used, and if you even have some ram, you shouldnt use it even for h3 since int8cr will be quite a lot better
>>109468013its not the same compression if it says int8, although i dont know why it also says nvfp4, maybe its mixed quants, maybe they forgot to rename properly
>dynamic vram and hip backend works properly with the new comfy-kitchen mergethank you to everyone who kept pestering comfyanon to take a look at it, i can finally gen my ex having dinner with me
>>109467728>>109467730>>109467642https://huggingface.co/Kijai/PrecompiledWheels/blob/main/sageattention-2.2.0-cp312-cp312-linux_x86_64.whl
>>109467702How to connect it bro ??
>>109468218>spoonfeeding
>>109466357biggest issue is that you can't use it to make an existing video longer without the audio getting cut off afaik.
>format spoken dialogue as instructed in minimax's prompting guide>english comes out like gibberishaudio isnt very great on this model huh
>>109468360it is kind of tricky though at least right now to build it on linux, i posted a solution to build from source in the next thread.
>>109468443I tend to have problems with audio too.I also think lot of the "free optimizations" people are running might also be playing a role.
>>109468360And the issue is the devs removed the wheel for linux and now it can't be built from source either due to newer version of gcc or cuda toolkit what ever but there are prebuilt on other sites. I keep it noted because I knew once the new big video models dropped anons would get stuck.
>>109468443I think audio is one of its best features, but as always, it's not an exact science and prompts, even when done exactly as in the guide, are not reliable. Hell I've had more luck with short, dumb prompts and it nailed the audio right away.
>>109467595>I used pizza tower as a style reference and this happenedMinimax H3 is starting to depress me, because I can no longer blame the model for not obtaining what I want. Anything that I don't get to see with AI now is a result of my own skill issues
>>109467620People who are surprised by trumps take on this forget the JD Vance is the only VP who had ever said the words "open source" in an official communication. This is one of the good things about the government being controlled by tech oligarchs
>>109467921>bf16There is exactly one consumer card that can run H3 at bf16 and you should still run it in int8 convrot because you save on time and the degradation is unnoticeable, including double pendulum effects, before 10 seconds into the video or so
>>109467620This only applies to US based models, not everything open source. Likely that there's going to be a ban/restriction on Chinese models.
bros, how many steps are you running H3 for?
>>109467960I need the prompt for this pls so I can make him follow a girl wearing a skimpy outfit through a city street thanks
>>109467794Just get it to spit out the prompt and then change what you want. Are you retarded? Why waste time handwriting boilerplate shit?>>109467978Probably an abliterated model like https://huggingface.co/huihui-ai/Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliterated-MTP-GGUF if you really need to do it on your own hardware. Otherwise, you can just fill in the gaps that the LLM shits out.
Can't find KJ's "MiniMax H3 Mem eff Sage attention" repo on github.pls halp.
>>109469204It's the kj nodes repo, just copy and replace it in the comy custom nodes folder
What's the strat to using H3 as an image editing model?
>>109469273make videos 0.1 second long?but you're right. I remember there was a paper where they compared models and it showed that video models had the best understanding of 3d. They basically became 3d models under the hood to be able to make sense of how objects move in space.
or you could probably make your video 1s long and set the framerate to 1 frame per second and then instead of save video have a save image. I'm just guessing here though.