Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109502333https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
Bleeds thread
minimax h3 video of the year award goes to #8
make some sleepy gens
Blessed Bake of Teamwork.
real thread here>>109503609>>109503609>>109503609
Will minimax even release their super secret magical upscaler that would solve the shit faces we have right now for H3?
>>109503683xD
>>109503684If we get stood up again like the chinese tend to i'm gonna declare jihad on the 100 acre wood
Why does debo the idiot keep trying? Even though he keeps failing?
>>109503683hahano
>>109503703its the pokemon spammer
>>109503683You fail at this every fucking day. Next to self soothe your autism you'll spam news>>109503703He has nothing else to live for
>>109503584>>109503551
>>109503706It's not working debo. You're not smart enough to change your posting style.
>>109503683
>>109503692I wouldn't be surprised, it wouldn't be the first time just the last little thing is withheld.
>>109503703it's not debo, it's whoever decided to shit on the thread this day
>>109503703It really would be better if he just died. Just committed suicide. The world will be a better place when he finally dies.
>>109503719I want him to live a long life but lose access to the internet.
Any way to make H3 work in other aspect ratios without simply genning in 16:9 then chopping the video up?
Animabros...
>>109503738
>>109503738I am using custom resolution just fine.
>>109503743What in the shit is that UI horror
>>109503743>making 1girls in ideogram
>>109503741>>109503743Thanks, I guess my early tests were just bad luck. Figured it was something forced by the training.
>>109503708>no quality degredation>sd1.5 lookinass
>>109503761>smear frame postingback to /a/
>>109503761Eh, I'll take the occasional mediocre in-between frame if it cuts my runtime in half.
praise be to op
>>109503739So anima is a meme then?
thanks for bakering
debo is the god of ldg
>>109503793benchod
The latest on /vp/:>>>/vp/59492200>>>/vp/59492256
>>109503793Always funny seeing you pathetic morons trying to test your VPN
>>109503752nta but Ideogram is still unmatched in photorealism
Hello?
>mfw Resource news08/08/2026>Kijai: MiniMax H3 Ref Lora Rank 256 bf16https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras>MiniMax H3 at native fp16 on pre-bf16 GPUs (V100 / Volta)https://github.com/Amduraznak/minimax-h3-fp16-fix>Cosmos3-Nano-WebUI: Self-hostable API + Web UI for Cosmos3-Nano quantized fp8 and nvfp4 checopointshttps://github.com/fengwang/Cosmos3-Nano-WebUI>R9700 AI Pro — ComfyUI / MiniMax-H3 speed patcheshttps://github.com/charlie12345/R9700AIProComfyUIPatch>MiniMax-H3-Pruned-GGUFhttps://huggingface.co/Abiray/MiniMax-H3-Pruned-GGUF08/07/2026>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely localhttps://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha>LIGHTX2V 4-step Turbo Minimax H3 lorahttps://huggingface.co/lightx2v/Minimax-h3-Turbo>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRAhttps://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA>Sage Ready: Local-only installer and readiness checker for SageAttentionhttps://github.com/CosmicFungi/Sage-Ready>Wan 2.2 Animate 2 14Bhttps://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUIhttps://github.com/NikoDemon80/ComfyUI-H3-Motion-Context>ComfyUI MiniMax H3 FirstBlockCachehttps://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache>KVAE: Family of Tokenizers for Multimodal Generative Modelshttps://github.com/kandinskylab/kvae>Energy-Guided Flow Matchinghttps://github.com/ysng123/EG-FM>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editinghttps://zzzmyyzeng.github.io/VideoArgus08/06/2026>Flash-VAED: Plug-and-Play VAE Decoders for Efficient VidGenhttps://github.com/Aoko955/Flash-VAED>(preview) MiniMax-H3 Turbo LoRAhttps://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora
>>109503813among other, not so favourable things
>mfw Research news08/08/2026>Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generationhttps://arxiv.org/abs/2608.04902>Coherence-Oriented Dream Scene Visualisationhttps://arxiv.org/abs/2608.05233>Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generationhttps://arxiv.org/abs/2608.05210>GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Modelshttps://arxiv.org/abs/2608.03083>IRIS: Visual-Semantic Binding for Forgery-Resistant Watermarking of Diffusion Imageshttps://arxiv.org/abs/2608.03539>Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generationhttps://arxiv.org/abs/2608.00663>A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrievalhttps://arxiv.org/abs/2608.05260>Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understandinghttps://zhangbo135.github.io/EviSelect>Progressive Decision-Making for Localizing Open-Ended AI-Generated Image Forgerieshttps://arxiv.org/abs/2607.29156>GeoMAR: Unleashing Geometrically Aligned Features for Masked Autoregressive Blind Face Restorationhttps://arxiv.org/abs/2608.03923>WorldClaw: Agentic 3D Open-World Generation at Scalehttps://arxiv.org/abs/2608.05248>UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Spacehttps://arxiv.org/abs/2608.03817>Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inferencehttps://arxiv.org/abs/2608.03867>Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Modelshttps://arxiv.org/abs/2608.03160>Attention is Case-Sensitivehttps://arxiv.org/abs/2608.03711>In-Context Collapse in Vision-Language Models and How to Mitigate it?https://arxiv.org/abs/2608.02830
>gen looks good>"oh but what about this added detail"at some point the prompt will just overloaded and stuff will start to fall out, but we haven't gotten there yet
>>109503820I hope the minimax imagegen model can bring the best of all worlds (non-autistic prompting, H3's uncensoredness, Krea 2 knowledge of pop culture in terms of characters and celebrities, Ideogram's photorealism, Flux 2 editing)
>>109503839>Krea 2 knowledge of pop culturethe video model seems way worse than krea2 when it comes to knowledge and characters
>>109503708miyako... my love
>>109503792true
>>109503819>>109503827thanks!
>>109503835why did you put two quotation marks next to each other?
>>109503846References work even with the text model you can do a style transfer too and it will know the character or pose, it's not as good as the other model but for my task it's been fine.>>109503858You're a waste of space and are gradually being gate kept out of this hobby, once next gen hits you'll be priced out.
>>109503863autist anon...The top one is their impression of the gen, the bottom is them saying that to themselves.Not that anon by the way
>>109503864yeah the reference is quite nice. if the image model can take references its not too bad then
>>109503860:)>>109503864nah
>>109503871what do you mean "their"? who is they? i only saw one person
turbo lora and sage attention are all I needimagine having more than that
>>109503739Shiina best girl...
>>109503873I wish krea2 could do that desu, even the style can be well preserved>>109503877I guess you're going to suck your way to a new gpu. Fair enough, you could try getting a job but....
>cucking yourself with snake oil instead of being patientcouldnt be me
So I actually read the schizo rentries and 99% of them is complete fabrications. care to explain author schizo?also who replaced the collage generator with lowfps vibe coded slop?
>>109503890based patient anon
>>109503879Saw? I'm talking about an internet post.
two pass workflowfirst pass>0.2mp>err_sde / beta57 15 steps>cope freesecond pass>RTX upscale to 0.8mp> eular / beta - 0.4 denoise - 4 steps>turbo lora + soltotal runtime: 4min16s on rtx 3090https://h.uguu.se/rUbOiawR.mp4
>>109503904wow, that's pretty good
>0.2mpguillotine
>Watashi, kirei? Translation: Am I pretty?https://files.catbox.moe/1evygl.mp4
>>109503904cool but sadly stereo is still a bit of a ways off, doesn't work well
>>109503890My brother. 0.8mp 15secs absolutely raw not even sage on a 3090. Feels good to be pure.
>>109503921honestly you can sage without worry, it's visually the same and way faster
>>109503924>honestly you can sage without worrty, it's visually the same and way fasterthis is true and if you're still skeptical read the sageattention2 paper. sageattention3 is lossy however
>>1095039040.5 denoise might be even betterhttps://n.uguu.se/TrkDYPcK.mp4
sage + mem eff sage + easy cache + turbo + torch compile + radial att + teacache + spectrum + mangekyo sharingan = izanagi
i cant stand that i need to reopen comfy every so often because it decided to just consume ram while doing nothing
>>109503541>KJ sage vs caches mixtureHuh. The cope caches mixture.... are pretty good? Both are cold start from boot up. 5090, ref model, 1MP
>>109503957where else would comfycloud get their hardware???
>>109503890>>109503921>when you're so afraid of breaking something by updating comfy you started to become patient.
>>109503961>tests it on tranimeevery time
repost.Sonnet II continues to be repainted.>> Anonymous 08/08/26(Sat)21:34:12 No.109503664▶>>109503648 (You)reminder this week's gensindia diss:https://files.catbox.moe/dioqb9.mp3sonnet IIIhttps://files.catbox.moe/sddmtd.mp3sonnet Ihttps://files.catbox.moe/w2ysvt.mp3
>>109503913https://files.catbox.moe/nk3udp.mp4
>>109503971i pee pee'd my trans panties. you owe me a new pair
>>109503971wow, eye contact. Don't ever get that shit.
>>109503975calm down and listen to thread music:>>109503970
>>109503962bro...i'm gonna cry
>>109503971im sorry, asians are just not creepy.
>>109503975what makes trans panties different from regular ones? the bulge pocket?
>>109503990I beg to differ
>>109503994has a dilator built in
>>109503971House 2 lookin good!
>>109503949yeah sage 3 isn't that interesting, even if it's fast, it's too lossy
Why do you guys hate sexy minimax gens so much?
>>109504000whats creepy about this?
>>109503970oh yeah and another india diss:https://files.catbox.moe/92cery.mp3
might be on to something with this 2pass wffirst passhttps://d.uguu.se/TfnamihY.mp4second passhttps://n.uguu.se/ppDynuxh.mp4I never had this gen look this smooth before.
Total /ldg/ Victory?
>>109504056>mp:^)>4*jazz music stopsHARAM!!!! ABSOLUTELY HA''M
>>109503739it really is a mess the more i use it. i wish he would've picked a better base. i hope we get a krea 'tune or something, it's hard to keep using when the prompt comprehension is so bad compared to krea.
>>109504064a small victory in combat does not mean the war is won
>>109504070are you disdaining the tillage of thy husbandry?
is pinned memory just flat out broken? it's jumping to OOM on 1s/0.1MP. I know just disable it, but I've got tonnes of ram just doing nothing and my gpu is only ever half full
>>109503924>>109503949guess I will taste the forbidden fruit. should I use the kjnode or just add it to the bat file?
>>109503739Who made this purposefully disingenuous "comparison"? kek
>>109503913>>109503971nice, was wondering when we'd finally start seeing the horror gens.
forsenhttps://files.catbox.moe/ca03s7.mp4
>>109504109Both do the same thing, kj is just more convenient since you can disable it without rebooting comfy.I personally have it on all the time so don't use the kijai node.
>>109504149sure feels like summer.
>>109504149how to get the layout/box:Use <Picture 1> as the reference image for Forsen. Use <Audio 1> as the voice reference for Forsen's dialogue.LAYOUT:The output must look like a live broadcast stream layout. In the top-right corner of the frame, there is a small, sharp, rectangular webcam box displaying <Picture 1> reacting live.SCENE & DIALOGUE:The background features an anime style Hatsune Miku walking in Tokyo, holding a green leek vegetable with her hand. Inside the square webcam box, the streamer Forsen maintains a completely relaxed, calm, and low-energy demeanor. Forsen moves their mouth with a casual, effortless talking cadence.Forsen says: "Wow what a great link, totally worth the 20 dollars, right chat? What a great use of your hard earned money."
Help. How can I generate creative thoughts?
>>109504161Read some books, be interested in things, stop to think, imagine, dream, hope
>>109504109i mean if you're lazy you can just add it to the bat file but kjnodes has a mem optimized one that you should probably use to savce vram
so things like <Picture 123> is going to always point to frame 123?
>>109503990The video is not meant to be scary, it's an artformhttps://files.catbox.moe/pemgwk.mp4
>>109504171I was waiting for the scary part....
>>109503904>>109503953Wait, you can use LTX upscale in Minimax now ?? Source ???
>>109504180it's not LTX, the WF is in the files, using turbo lora as upscaler
>>109504189No i mean Minimax dont use two pass gen type in the first place. Is that two pass gen ?
>>109504161if you can't do it personally, you probably can with an LLM
>>109504202>Minimax dont use two pass gensays who? is it a rule? did I break the law?
>>109504161scroll civitai redit may be porn but it will spark ideas
>>109504202nta Is just akin to a hi-res pass just with video gens. Is not really specific to LTX, you can do with any of the models. LTX just had it's own infrastructure around it
>>109504209Youre the first guy who did it. Share your tech to everyone
>>109504217it's in the workflow!!!
>>109504176same i expected someone to be behind her, would be a perfect use of AI
hahahafirst time trying a non english voice clone gen.https://files.catbox.moe/64h6np.mp4
>>109504149Kek, the voice is so realistic I thought it was really him commentating an AI game sarcastically for a secondhttps://files.catbox.moe/wkf1rj.mp4
of course it's hitlernext it'll be floyd
>mfw it consistently replaces physique, clothing and hair, but not the fucking face, despite explicit, correctly formatted instructions
What vm are you guys using for cumfart? Just docker right?
>>109504258maybe it'll be miku and hitler dancing together
>>109504176>>109504224Anon I'm just testing how well it can do 1girl vlogs for gooner purposes, though it would possibly would need a LoRA to make her masturbate and say dirty things
>>109504236just plug english in an english to german translator, use "says in german", plug in a downfall hitler voice sample in ref audio 0, and voila:Use <Picture 1> as the reference image for Adolf. Use <Audio 1> as the voice reference for Adolf's dialogue.https://files.catbox.moe/6si4d6.mp4
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/tree/mainAnyone tried Step600 ??How was it ?
kinohttps://files.catbox.moe/na4mn9.mp4
>>109504310>4looks like it's not an mp3. Maybe it's some kind of mistake, I'll allow you to resubmit.
Should there be any quality difference between int8 and nvfp4 CLIPs?
>>109504289step 600 ema is the best one to use currently.
What's the best image edit model for anime right now? I need to start building last frames.
Using H3, I am going to personally revive the 90s late night sleazy thriller movies. It will be a new golden age of boobs and violence!
>>109504324Not that I've seen. I now exclusively use nvfp4.
I hope you boy's are behaving yourselves in here.
>>109504385bomboclat
>>109504330Whenever H3 releases their model...
>>109504330Anima.Ignore the schizo.
>>109504400I use anima for generating the first frames already, it's great for that, but as far as I know there's no good tune for image editing.
>>109504414Then use Qwen Image Edit 2511
>>109504330local edit is absolute garbage, just like video was until h3. people will cope and tell you klein or qwen but they're just freetarded.
asuka test (ref)https://files.catbox.moe/m6l8zp.mp4
>>109504417Thanks, I'll try>>109504438I'll see I guess
>>109504444checked
time flies so fast when genning its not even funny
>>109504262Well its a quirk from the model, at some point people will figure out how to train character loras. At this point after autistically tadwrangling it for 5 days i finally got it to be able to actually use the video reference for motion and initial positions. I still get hit by generic face and some anatomically horror beyond man comprehension here and there even with 9 reference pics from different angles so at this point i made peace with it and just gonna wait and focus on preparing the data sets to train the loras when there enough research about how to do them properly If you are extremely autistic remember that video are just a gorillion images so you can just use klein to faceswap every frame with some masking nodes
>>109504438Sadly true. Also lacking a good music model.
sometimes i get "lazy" gens that only animate the specific most talked-about part of the i2v prompt, especially when the input image is an art style you wouldn't often see animated (hand-drawn, high quality art etc that it's rarely practical to animate), anyone found reliable workarounds for this? experimenting with "the whole image is animated" etc at the moment but maybe it's better to just explicitly spell out lots and lots of parts' motions and eventually it'll be animating half the image and decide to just animate the lot?>>109504330>>109504414try qwen image edit or flux klein 9b, but yes, >>109504438 is basically right, edit's very hard. if it's SFW you can use api models of course, they have basically perfect edit, eg gpt-image-2
>>109503713The greatest output to ever come from LDG
So, Today summary is only ComfyUI 0.31 which is support Kijai int8 VAE and this >>109504289 huh
>>109504480>anyone found reliable workarounds for this?Other than a reroll and a prayer? No, I fucking wish.
pretty impressive considering all I plugged in was baka shinji, 2 seconds only:https://files.catbox.moe/p5vegl.mp4
sorry guys i was too busy grilling up some steaks. kino will be generated soon
There's a bug :(
>>109504521your wires make me sad. the bug is between the user and the screen
>>109504523i think you mean between the chair and keyboard retard-chama
>>109504523try it, minimax doesn't work on a ksampler if you put>add_noise disable
>animeArtDiffusionXL_alpha3
how come describing facial features is shit in every single AI modelYou'd think being able to describe the shape of the eyes, nose, mouth and face would matter a bit more
Did they messed something up again because my gen got slower with ComfyUI 0.31
probably
>>109504257>east asian feetthe best there is
Based!>YES WE WILL KEEP OPEN UNTIL AGI ARRIVES.https://xcancel.com/MiniMax_AI/status/2086253065657790895
>>109504438why can't h3 be used as an image edit model?
>>109504585They've already said they're going to make an image model and edit model.
ace step in the oven. is everyone else an indian in a cybercafe with only free horrible headphones available?
>>109504580agi doesn't even mean anything at this point
>>109504480>>109504438Yeah, qwen image edit kinda sucksTrying to create anything remotely pornographic just doesn't work, and I can't find any finetunes for it. What a shame.
>>109504571Nevermind it was the turbol ora fault. larryvrh fuck something up again >H3TURBO fwd full/bypass] call#1 qkv_proj.forward_owner=BypassForwardHook (BypassForwardHook => lora ACTIVE; else => BASE ONLY!) is_injected=True timestep=1000.00 video_rms=1.0000 audio_rms=1.0048 dtype=torch.bfloat16
>>109504588Someone should get ahead of them and make it.
>>109504580we are literally living in the best timethe time between when ai is first starting out and we still have full control, and the great filter in the next 5-10 years
>>109504585you can, set length at 5 and extract the first frame. voila.
>>109504611>great filterI have faith, we'll wing it like we did the nuke filter
>>109504496rei test:https://files.catbox.moe/c326bl.mp4
>>109504623saving this for your notes alone. thanks nerd
>>109504623there we go, close shot of Rei worked nice.Use <Picture 1> as the reference image for Rei. Use <Audio 1> as the voice reference for Rei's dialogue.LAYOUT:Close shot of Rei. Rei is standing on a sunny beach in Tokyo. Rei says in Japanese "(put english to japanese test text here)".https://files.catbox.moe/vtcrlc.mp4
>>109504257>>109504574If only it could always spell right kek, I would regen but then I have to wait another 10 minshttps://files.catbox.moe/2o2uxa.mp4
I wonder, if there was an extension to batch process gens with a script, in theory you could generate a show while afk, even.>sequence of 15s clips>stitch them all together
>>109504580>until AGIDo they mean that they commit to releasing their hypothetical AGI model, or that they'll stop sharing as soon as they reach it?
https://files.catbox.moe/5xtze0.mp4
Use <Picture 1> as the reference image for Rei. Use <Audio 1> as the voice reference for Rei's dialogue.LAYOUT:Close shot of Rei. Rei is standing on a sunny beach in Tokyo. Rei says in Japanese "Totemo sutekina katada to omoimasu. Motto ohanashi shimashou!". the girl from <Picture 2> walks in from the right and says "Miku, dayooo".https://files.catbox.moe/dy4wax.mp4
alright, it's high time i stopped lurking and posted a vidhttps://files.catbox.moe/1rldy7.mp4
>>109504663Especially since you can use fl or ref and just indicate the last frame of last video is first frame of new video
>>109504672long shower before + bag over head during
>>109504679light chuckle
>>109504679lole
>>109504392https://files.catbox.moe/mw4ybn.mp4You shall not pass.
>>109504699>claude, write me an edgy prompt.
a big anime for you:https://files.catbox.moe/xyrqhm.mp4
H3 can do video editing.
>>109504723which is which
>>109504615Why not set the length at 1 frame?
>>109504737nta but you can't it's just how the model works
>>109504723You missed the shadow
>It doesn't understand "bandaid underwear">Show it a reference>It gets it the very next genman I love living in the futurewell, well worth 25% longer gens over the first/last frame model
>>109504723It can't remove/swap minor details at specific timestamps.
>>109504664The latter
>>109504611no, it's terrible time
>>109504743>ntawhy do redditors always say this? just answer his question without saying that, nigga
>>109504754Oh, and I tried describing it about 30 different ways, exact placement, super literal descriptions, nothing worked. A reference instantly solved it.
What should I generate?
>>109504580>until AGItranslation: until it's competitive with seedance, and I have no doubt minimax h4 will be competitive, so we'll never get a model ever again from those fags
>>109504767>redditors>not that anonhmm
>>109504672pajeet
>>109504672this isn't ai
>>109504785what?
>>109504580china is so fucking cool seriously, do you see a single western company screenshotting a "The Rock" reaction image and posting it officially on twitter??
>>109504792correct. both the webm and the catbox are real.
>>109504779If the "Seedance is 200b" rumors are true, they might just have to train a Large version of H3 and put it behind API
>>109504679
>>>/wsg/6210720Ref model with an audio clip and just prompting something like "Dance to the music of <Audio 1>" works pretty well.
>>109504653Total tomboy enjoyer victory
>>109504580I still remember back when anons were talking about how Kling or Minimax will never be open-source,, all we had was Wan
>>109504679this is why I pay for the internet!
>All development slowed down todayGuess we all running out of semen huh
>>109504793you're the redditor if you think "nta" is a reddit thing, it has a different meaning on 4chan, newfren
>>109504723Yeah but it's a major pain in the ass to get it working though. You have to write a hugeass prompt depending on the complexity, and get a good enough video where the model can figure out where is what.I could only do the picrel (out of the MGS3 big boss gif) with an LLM-written prompt based on the official docs
>>109504816>I still remember back when anons were talking about how Kling or Minimax will never be open-sourceI said that, I never expected Minimax to make such a move, not that I'm complaining though!
Ugh. Really hate it when i got low framerate animation in Anime style
my pc just segfaulted right at the end of a 20 min gen...
>>109504672gross, that’s indian
>>109504830>you're the redditor if you think "nta" is a reddit thing, it has a different meaning on 4chan, newfrenno, you're the redditor if you feel the need to deanonymize yourself just to give a simple answer to a question. fact.
>>109504815It's all that gay stuff. Just admit gayness or get a real woman.
>>109504580>The answer is YES: we plan to open-source a unified text-to-image and general image-editing model derived directly from the H3 lineage. It is currently in post-training refinement.HOLY SHIT WERE SO FUCKING BACK WTF
>>109504841https://civitai.red/models/2843112/astrowitch-cinematic-comic-style-lora-by-astroburner-ai?modelVersionId=3209714>Please pay to access this model goyFuck off
Why did you guys write off wan fun-VACE off so quickly? I think you are all overrating minmax and underrating wan's true capabilities.
>>109504828This is a civilization destroying technologyUltimate dopamine machinesPeople are probably getting addicted to genningAs more barriers are removed it will become worse
>>109504653idk, doesn't seem to have the same crazy errors as krea does.I'm not too happy with the feet, but they are not quite goofy Gnome feet.
>>109504858dont make me feel old you son of a bitch
>>109504580surely a CEO would never lie
>>109504854oh ok, sorry for misinterpreting your point, i disagree with it but you're entitled to your opinion
>>109504856nahh, tomboys are the ultimate expression of femininity, if a woman is still beautiful with short hair that means she's naturally way more feminine than your average woman who has to get her hair long to be attractive
>make a unique workflow with help from claude to solve the errors I'm getting>spend like a week on it and finalize it>works amazing>update comfy>workflow still works but now has blur in all of the gensNIGGERS
>>109504870he doesn't need to lie here, his post is so vague it can be interpreted in any way he wants, "AGI" can mean anything
>>109504871thank you
i click on every video because i know anon took a lot of time to plan it and generate the output
>>109504857>>109504580so we'll probably get Minimax edit before Krea 2 edit huh? interesting
>>109504887I love you
>>109504679fucking slop so is everything here.
If Minimax always force you to get low FPS on anime, theres no reason to have 24fps genning render
>>109504860its all fucking slop all of it why would you care? every krea lora is fucking slop everything is slop all fucking slop.
>>109504909>>109504895>>>/v/
i deleted 2 TB of slop and i feel better, i am free.
>>109504909Don't fight the slop..embrace the slop. The slop is your friend
I hope the Minimax devs use what they already have for H3 to make an audiogen model too. Shouldn't be expensive since they would essentially just have to fine-tune it to generate long audio. One audio model that does everything: TTS, sound effects, music. Bonus points if you are not forced to autistically add key and bpm into the prompt like AceStep requires.
>>109504919damn, anima krea and h3 really take up that much space?
be not afraid (snafu)https://youtu.be/JvjxOLRNLeIhttps://suno.com/s/zMaaz8vjHjZ1Nam1
Seems like the hype winded down, threads were def. faster the last few days
>>109504931pretty harsh on the ears no?
>>109504929No Anima, Krea, H3, XL. I only have left my one and only friend SD 1.5
>>1095049461.5 is just 1.4 but slopped
>read h3 loras comments on civitai>most of them is just users complaining that they get body horrors and they don't work
>>109504929
the scalping solution:https://files.catbox.moe/s6m3s9.mp4
>>109504953loras are gonna be hard to train, they only gave us the guidance distilled model
INT8 Video VAE saves gen time by 30 secsNice
>>109504965also fuckstarts the quality
>>109504893I would not mind for them to tease us a bit with some image showcases, if they said they almost finished it then it probably looks good already
>>109504944cum get thinned out bro. Need days to recover
>>109504876>>make a unique workflow with help from claude to solve the errors I'm gettingwhat errors could you possibly be talking about
>>109504965i would not do that if i were you.
>>109504873It's mere homosexuality.Like having women on scotus or whatever. Mere homosexual tendencies.Real men put women under strict rules. Certainly they don't vote.
>>109504953Keep in mind that users on civitai are extra retarded.
>>109504992lmao what
>>109505006>anon always thinking about queersyou're gay
Just one more genone more genone moreonemore
>>109505006you also hate women another red flag or indication that you are gay.
>>109504960better scalping solutionhttps://files.catbox.moe/5672u9.mp4
the `text generate` node got a bit more retarded with comfy update. impressive
>>109504925You can audio gen with H3 by generating a video with the minimal resolution (32x32) and a long length (~1 min).>>>/wsg/6209660
>>109505019Never sleep Only gen
>>109505009 If your lora cant be used by retards then its badly trained
>>109505079only a retard would say that
>>109505040>swap to the next `text generate` node available in cumfart>same settings, thinking enabled>doesnt start off mid-sentence, doesnt copypaste the same prompt i put in 3 times, doesnt cuck the gen>everything just workssomeone on the inside is trying to sabotage comfy
>>109505019Just one more gen and I'm going to bed. One more.No the preview of that one's not good enough. Cancelled.Just one more gen and I'm going to bed. One more.
>>109505036That scalper shouldn't have killed his dog!
>>109504619>we'll wing itthere's no "we" here. we're going to die to a "black hole in the garage" event where someone unleashes a paperclip maximizer and everyone gets turned into a mcdonalds hamburger>>109504679kek her facve in the last frame
R2V turbo lora when? Please, I need it.
>>109505024Yeah, everyone hated women before women had the vote, just like in the reddit moooovie slop.
>>109505107The power behind r2v is incredible. Imagine when we get an upscaler, loras, any official updates from Minimax - it's gonna be impossible to do anything but sit here and gen
>>109505122sadly thats not gonna happen, the team behind h3 will be hired for a private company and those models will never be released, just like happened with z-edit
>not a meme
I need to a sprite sheet for a project, I dont care so much about the style or being blown away by how great the visuals are. I just want it to be functional/consistent between sprites. What would be the best model for this?
>>109505122>it's gonna be impossible to do anything but sit here and genim already getting bored
>official updates from Minimax>
>>109505137>z-editZ-image edit status??https://files.catbox.moe/sf09k4.webm
>your cpu>your gpu>your ram>how youre feeling about itgo
>>109505148you're gonna get maxxed out by Tyrone in jail, pedo
>>109504887I use JavaScript to delete all posts with videos from the rendered HTML on 4chan and Reddit via tampermonkey.It's the closest thing to a genocide of these people, and it's the best experience.
>>109505122h3 already hit plateau, anything complex will filter 90% of lazy AI jeets, and if you hadn't already noticed, anons like >>109505157 got out of ideas and they just repost their unoriginal funny slop, oh its that another Seinfeld but they don't look anything like Seinfeld h3 slop? or is that another the big bang theory h3 clip
>>109505157can he do the heart rip one?
I continue to do repaint work with ace step 1.5 xl base, on Shakespeare Sonnet II
that's one gigantic ass (adult woman:-1) workflow kek
>>109505112no you are gay. you do not understand human biology.
>>109505178the reference model can make anything look 1:1, there are no limits with the ref model. image/voice/whatever.https://files.catbox.moe/na4mn9.mp4
>>109505148Total Janny Death
>>109505143good.
>>109505195all deviation from the Eve ideal are faggotry. I'll die on this hill bring it. If you don't want Eve, YOUR GAY.Eve: basically Venus.
>>109505197wow yeah, it looks amazing, gj anon
>>109505157https://www.youtube.com/watch?v=0oMEuyhBkRo
>>109505201Nah, pretty sure they deleted it themself. The workflow from the catbox has some leftover prompts in it lmao.
>>109505208yes it's a 0.3mp test render ofc it wont be high fidelity
>>109505208>the ai learned to have the same blur as moviesidk man
>>109505178not my gen i just saved it in case i need to post one that isnt mine for a pointless post
>>109505215>they>themselfwhat do you mean? did multiple people type out that single post?
>>109504672some poltards actually think this shit is real>>>540459779looks like my work today is done
and of course it doesn't resemble the German prophet, but that's clearly a feature to avoid profaning him by creation of a graven image.
>>109505223uh oh, esl melty
>>109505208name another model that can make AVGN like this.https://files.catbox.moe/xqsl9a.mp4
>>109505228idgi
>>109505197>https://files.catbox.moe/na4mn9.mp4KEKKKKDD
>>109505157Two pcs:96gb ram50909950x3d96gb50905950x3dFeels pretty good.
>>109505148those bypassed nodes aren't chinese cartoons..
>"she is talking animatedly">every seed is just two hands raised and mouth open
>>109505238nice
>>109505254>one for waifu>one for tiger mom life coach>one for vibecodez>one for gaymer>one for blender>one for music diffusion>one for image diffusion>one for video diffusionyou don't have enough.
the harsh true is that h3 turbo loras sucks, they only look good in 2d animations or close-up shots, any complex motions looks like shit or stiff, light2xv has better motion of they have been using the same synthetic dataset since their first turbo lora, so it changes the face of the subjects
okay. NOW the scalpers are finished.https://files.catbox.moe/tmdb5o.mp4
>>109505270There aren't any complete turbo loras yet thoughbeit
>>109505238the problem with these more static gen styles is it's not exactly pushing the capabilities of H3, it's the kind of thing LTX could do - that is not to say LTX would be as good but it could do the voice and the emotion - you could literally do the soundtrack you wanted in a t2s model and that would be even better.
>>109505270anon, wait for them to finish the cook, it's only been a week
>>109504760Well, I did try, but if the cyclist do not stop, the sign must insist.
I haven't been genning minmax all day. I've been studying and brainstorming the capabilities of R2V.I'm not bothering to post any of my test gens here, but I am making progress. ChatGPT 5.6 Sol has been very helpful through this.
does this suffer the same problem as SCAIL where the character swapped matches the body of the target video?
>>109505297kek
>>109505288The reference control pushes it leagues beyond ltx, even with more static gens.With ltx you could generate a start frame with an image model and then pray it works as you wanted it, but with H3 you can have whatever characters doing whatever whenever with whatever audio thanks to references.
>>109505279no one fucking cares man, get some imagination, like really get into knowing how to imagine a world and create it, start simple. let the model give you ideas, don't use speech prompts just say "2 men talk about thing" and let the model fill it in, it will be mostly gibberish but it will include some English about that thing but the simlish is what inspires.
>>109505306no i don't think it does but that method is hard to prompt and takes longer to process.
>>109505297ngl this goes hard, like the sign won't stop following you unless you fucking stop
>>109505311ltx could do references too
>turbo lora>shift enabled>spectrum enabled>dont change steps from 20shit whoops
>>109505314anon, we are full throttle into retard/autistic h3 cringe slop, their creativity and unfunny shit just replicates how they are in real life
spectrum was absolutely shit in WAN. why should i trust it'd not be shit in H3
>>109505328Oh huh. I spent like 3 days using it, generated nearly a thousand videos, and only liked one of them so I stopped. I genuinely had no idea it could do references.
>>109505338and your constant bitching replicates how you are in real life?
>>109505250add english subtitle
>>109505142anyone? :(
>>109505311That is a different thing. If you are basically doing a guy sitting still from that source image then it's not really going to be different unless you start do cutscenes which is kind of weird because why would a guy reviewing shit be doing editing? Yes I realise he is a parody e-celeb doing "reviews" for comedy value. Not that I have ever watched him anyway so I don't' know if his delivery style sounds so amateur as usual e-celebs have a better vocal delivery and are more professional - he sounds like he has never done anything in front of a camera before.
>>109505360krea2
>>109505346you tell me
>>109505367ok ty
>>109505360>>109505142probably gpt image if its sfw. if its nsfw i dunno
bakin
>>109505342it kind of works like the FL2V model. you put your characters into the starting frames and you can also put audio at the start as well. then you start the prompt for the scene you want
>>109505142>>109505360Unfortunately it's SaaS / not local, but check this:https://www.autosprite.io/
>>109505342it really is somewhat similar to what H3 is doing but obv it isn't as good because the LTX model isn't as good. But LTX used within it's limits works pretty well. Use a short audio ref (say a voce) and it will replicate that voice just like you can in H3 with the talking parts in the prompt. I used to use an enhanced prompt so that the speech was varied around a subject matter for them to talk about.
>>109505389make sure to put my gens in the fagollage or else i WILL sperg out
>make sure to put my gens in the fagollage or else i WILL sperg out
>>109505402how I'm supposed to know what gens are yours though? kek
>poorfags unable to animate their favourite jaks
>favourite
Minimax doesn't know teto so I had to use the reference modelhttps://files.catbox.moe/aswsua.mp4
https://www.reddit.com/r/StableDiffusion/comments/1vj5scs/deroping_minimax_h3_fast_motion_to_reduce/https://github.com/matlowai/ComfyUI-MAINodes>H3 smears bursty motion: backflips, fast sword arcs, whip-fast reversals. The cause is structural. One latent token spans four pixel frames, and at high motion speed those four frames need four distinct poses that a single token can't hold. Re-denoising the affected region doesn't help, because the missing poses were never generated in the first place.>This pipeline works around that at inference time. It re-generates the clip as a slowed-down version of itself, seeded from the original. Frames where motion is too fast get held (repeated) so the model has more temporal room, the result is generated video-to-video from that retimed init at partial denoise, and the original frame rate is recovered afterward by dropping the held frames. The oracle that decides where to slow down reads the clip's own latent. No extra model, no training.
tv is savedhttps://files.catbox.moe/eksdh5.mp4
>>109505420mine are the good ones obviously
>>109505402im not actually bakin
>>109505444people cried about the final season so much I decided to never watch this show, thanks for beta testing this for me fags!
>>109505443>oracledemon
>>109505397>>109505379hmm if i cant get usable results with local models ill give this a go, thanks
>>109505444checked 10/10thanks for keeping the file lmao
>>109505443sounds amazing and also that the computer has to do some extra work
Blessed thread of frenship
>>109505444>>109505470unironically the first video gen in the world I've saved afaik
>>109505449no one wants to run the malware i take it? cause i'm not bloody trusting it.
>>109505484i refuse to make these threads on the premise of i dont care i got other shit to worry about
>>109505489>>109505489
>>109505484ask claude if the code is sus or not, that's all
>>109505492that benchod can't even go through a 3kb file at high effort without telling you to upgrade
>>109505492still won't trust it, it could change any time, id rather do something like that myself, it could be knocked up pretty fast i'm just too lazy.
>>109505500>it could change any timewhat
>>109505500i mean i had similar bash scripts using ffmpeg to make a canvas so that different size videos could be merged. similar can be done with the collage. I do not trust code i can't read sorry.
>>109505505famous last words people. >No how could they have done this?!?
>>109505505if i do not understand the code i won't use it.
>>109505530they couldn't, it has no update URL and the one dependency you can just copy paste into the script
MOOOOOOOOOOOOOOOOOOOOOODS
I'm trying MinMax H3 for the first time, it takes 15 minutes on a 4060ti for a 10 second video at 480p, at default settings (slow but relatively acceptable although not at WAN levels with 4-step LORAs). I wonder if there are already 4-step LORAs for MinMax.https://files.catbox.moe/jdugu9.mp4
>>109507221there are scuffed ones but we are waiting for official ones
Playing around with the T2V H3 model now for a bit and really struggling to get truly dark scenes. Even if I prompt stuff like "dark room, darkness, dimly lit, video shows a very dark room, chiaroscuro" etc. it always wants to add studio lights on the people in the shot. Any anon with tips here?