Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109499734 >>109500977https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
>mfw Resource news08/08/2026>Kijai: MiniMax H3 Ref Lora Rank 256 bf16https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras>MiniMax H3 at native fp16 on pre-bf16 GPUs (V100 / Volta)https://github.com/Amduraznak/minimax-h3-fp16-fix>Cosmos3-Nano-WebUI: Self-hostable API + Web UI for Cosmos3-Nano quantized fp8 and nvfp4 checopointshttps://github.com/fengwang/Cosmos3-Nano-WebUI>R9700 AI Pro — ComfyUI / MiniMax-H3 speed patcheshttps://github.com/charlie12345/R9700AIProComfyUIPatch>MiniMax-H3-Pruned-GGUFhttps://huggingface.co/Abiray/MiniMax-H3-Pruned-GGUF08/07/2026>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely localhttps://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha>LIGHTX2V 4-step Turbo Minimax H3 lorahttps://huggingface.co/lightx2v/Minimax-h3-Turbo>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRAhttps://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA>Sage Ready: Local-only installer and readiness checker for SageAttentionhttps://github.com/CosmicFungi/Sage-Ready>Wan 2.2 Animate 2 14Bhttps://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUIhttps://github.com/NikoDemon80/ComfyUI-H3-Motion-Context>ComfyUI MiniMax H3 FirstBlockCachehttps://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache>KVAE: Family of Tokenizers for Multimodal Generative Modelshttps://github.com/kandinskylab/kvae>Energy-Guided Flow Matchinghttps://github.com/ysng123/EG-FM>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editinghttps://zzzmyyzeng.github.io/VideoArgus08/06/2026>Flash-VAED: Plug-and-Play VAE Decoders for Efficient VidGenhttps://github.com/Aoko955/Flash-VAED>(preview) MiniMax-H3 Turbo LoRAhttps://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora
>nigbo malware
>>109502332>not having an LLM to handle the prompt in real time
>mfw Research news08/08/2026>Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generationhttps://arxiv.org/abs/2608.04902>Coherence-Oriented Dream Scene Visualisationhttps://arxiv.org/abs/2608.05233>Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generationhttps://arxiv.org/abs/2608.05210>GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Modelshttps://arxiv.org/abs/2608.03083>IRIS: Visual-Semantic Binding for Forgery-Resistant Watermarking of Diffusion Imageshttps://arxiv.org/abs/2608.03539>Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generationhttps://arxiv.org/abs/2608.00663>A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrievalhttps://arxiv.org/abs/2608.05260>Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understandinghttps://zhangbo135.github.io/EviSelect>Progressive Decision-Making for Localizing Open-Ended AI-Generated Image Forgerieshttps://arxiv.org/abs/2607.29156>GeoMAR: Unleashing Geometrically Aligned Features for Masked Autoregressive Blind Face Restorationhttps://arxiv.org/abs/2608.03923>WorldClaw: Agentic 3D Open-World Generation at Scalehttps://arxiv.org/abs/2608.05248>UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Spacehttps://arxiv.org/abs/2608.03817>Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inferencehttps://arxiv.org/abs/2608.03867>Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Modelshttps://arxiv.org/abs/2608.03160>Attention is Case-Sensitivehttps://arxiv.org/abs/2608.03711>In-Context Collapse in Vision-Language Models and How to Mitigate it?https://arxiv.org/abs/2608.02830
>>109502333blessed triplets of frens
>>109502333thanks for the bake
hmm, now everyone is using gemma for prompting minimax h3 while no one was doing it for ideogram despite similar rigid prompt formatting requirementsinteresting isn't it?
>>109502354>how is he so fucking fast with his garbage news?probably a bot
breast thred of goonship
>>109502335i do not wish to use a cloud service>>109502334way ahead of you but what about the other one? im horrible at describing abstract shit
>>109502367ideogram kept saying what I was doing was haram and it kept sending my location to qatar
The model is capable of doing really crazy stuff with the camera, full POV and all that.Now I'm wondering has anyone tried generating (or continuing) an SBS video with H3 yet? I'll be doing it tomorrow for science.
>>109502374Use gemma locally for that shit
>>109502367It's just an effort to reward ratio, and the main problem with ideogram was the cuckboxes you were forced to draw for every gen
>>109502378I'm going to generate an H3H3 video with H3.
>>109502377>minimum number of bboxes: 5just add this to your prompt generator instruction and it generates anything without triggering filter
>>109502383any node recommendations for this so i dont have to install another llm toolkit or anything? i wont ask you for specifically amd solutions but ideally if you have one that would be handy
>>109502390ideogram doesn't generate hot trannies in my area. not interested.
https://files.catbox.moe/lreg9t.mp4good shit
>>109502333Damn, it definitely helps though. Like night and day difference. I'll try it with the Turbo LoRA workflow nowBaselinehttps://files.catbox.moe/8rzon3.mp4Fixhttps://files.catbox.moe/128mbo.mp4Fix, slo-mohttps://files.catbox.moe/a0frh1.mp4
>>109502398real shit fr fr no cap
>tranimeI'm out
>>109502398oh my god i've not heard that audio in ages
>>109502398
>>109502402Mean to quote>>109502328
>>109502402That seems smart. Gen bunch of shit and second pass only the stuff you like
So /ldg/, are you guys more of a [keyframe completion] community? or a [reference generation] community?
>>109502402>>109502417>2 mn to 20 mn
damn ok. turbo lora is fucking retarded with ref model. good to know.
>>109502408https://files.catbox.moe/rda7zi.mp4feast your ears
>>109502432>model drops>takes a long enough time to make one gen even on high end that people already make several turbo loras>several custom nodes to speed it up as well including various attention patches>just to make them gen faster at the cost of quality>the one thing to fix the quality also fucktouples the gen time>we will soon see workflows so fucked up from the speed that the base gen will be 20 seconds and the second pass will be 30 minutes
with the ref model being so good with high-res sources, we might actually have a workflow somewhere in the future where the gens are extended beyond 15s accurately in a single run and stitched.I'd try to make one myself but I'm lazy and retarded
So this is the power of the turbo lora.
>>109502438this is what glimpsing the infinite feels like
what did bro do?
>>109502448>I'd try to make one myself but I'm lazy and retardedyou already made it just by having the idea, now paste your own 4chan post into an LLM
>>109502435Works for me, use 600ema version, 8 steps euler/beta
>>109502450kino. what the hell
>>109502456see >>109502449I used those exact settings. just made it really retarded. The output looks really nice for 0.5mp tho. no artifacts.
Anyone tried https://github.com/matlowai/ComfyUI-MAINodes?Good or bad for face for ants syndrome?
>>109502450>not the millionth shitty meme about fent, floyd, niggers or jewsKINO
im not getting good results at all with the new 600 turbo lora, everything is all artifacty and grainy. doing 0.5mp. did i set this up wrong somehow?
>>109502470Purchase an advertisement
>>109502484Reduce strenght to 0.9
once again trying cope sol-attn
>>109502461I think prompting matters too, I had bad quality gens on H3 when the prompt isn't coherent, similar to LTX bad experiences when prompting something the model/text encoder doesn't understand right, and sampling is fighting with itself turning the results into mushy garbage
syntax highlighting added for MY prompt enhancer ;)
>>109502495well, it's better than spectrum, but... still shit.
>>109502485Get some publicity material
>>109502499*claude prompt enhancer
>>109502498I agree that the prompt is important. This isn't a prompting issue tho, the gen was fine before trying it with turbo.
>>109502501>it still causes morphing.why do people use this shit?
>>109502484remove external turbo sampler node, select eulerthat node is ass, designed for 4steps
>>109502527to stop coping one needs rtx blackwell 6000 pro. EasyCache, spectrum, sol-attn, turbo lora - everything's shit.
>>109502538I don't have that amount of disposable income.
>>109502540sir, comfy cloud is only $20 away
>>109502538The only two optimizations that aren't too insane : - sage attention 2, no perceptible change from my tests, and nice boost- int8 convrot model (same, no change for h3) using the pruned model
>24gb vramlets coping use comfycloud to access enterprise hardware for cheap. vidgen isn't mean for your shitty 2022-era hardware
>>109502540Then pay with your time
>>109502538I'm considering it, but very reluctantly. They can't get more expensive than the current $13k, right?
>>109502544>>109502541and how I should going gen my cunnies using copper cloud sirs?
>>109502549holy shit, that's what they're at now? i remember when they were $8k
>>109502538So how quick does it gen a 1mp 10s vid?
>>109502549They probably will if nvidia next release cycle is to be trusted : - blackwell gaming "SUPER" refresh for early 2027- actual next gen 2028-> means these 96GB cards will be in demand at least until then.
>>109502551fearlessly
>>109502549You mean barely $20K at the end of the year
at least sol doesn't seem to fuck with prompt adherence
>>109502549at my place they already cost like 20 grand
>>109502499You can do the exact same thing with LM Studio. And it's got way more features.
I've been using Qwen3.5 9b to generate prompts for H3. How does it compare to gemma? Worth downloading gemma?
>>109502571gemma is more creative
Buying a GPU right now is retarded. In a year or two, the AI bubble will burst, GPU prices will come down and Nvidia will shift its focus back toward the consumer market. We'll probably start seeing enthusiast AI hardware with much larger amounts of VRAM. I'm talking things like 64GB for around $1k
>>109502571Qwen sucks. Gemma is better. I get best results with 26B A4B heretic.
>>109502537still came out just as grainy, fuakguess i'll just wait for a better turbo to come out later
>>109502574lol
>>109502449for reference, same seed>no turbo>sol-attn, err_sde / beta57 15step>0.5mp>395sec (6min35sec)>rtx 3090
That's good bait
I remember my first Turbo lora
>>109502574this is a extremely dangerous level of copium to ingest. you should consult a medical professional immediately
>>109502432>>109502441Not so fasthttps://matlowai.github.io/ComfyUI-MAINodes/#ladderThat middle gen only took 2.2 mins for the 2nd pass and looks better than raw Turbo and raw baseline
>>109502574i listened to people like this about ram, and look where that got me
>>109502574>I'm talking things like 64GB for around $1kThis should have happened long, long years ago. Dram was cheap as dirt yet nvidia only supplied 8GB for gamers, and that in the age of quickly rising 4K, forcing game devs to use retarted techniques to have the textures looking okay.
>>109502578You need 22k in computers to make a good prompt
The reference model is genuinely magic. I can't believe some of the shit I've managed to do with it via motion transfer. Literal holodeck tier shit.
>>109502574yeah yeah, bubble will burst, we will find our tech jobs again, gpus and ram will cost less. simply won't haplpen.
>>109502589The sweet...sweet turbo lora...I ALWAYS HATE IT IT!
>>109502607kek
>>109502601>deepsuckkek
>>109502604Does it really do motion transfer without fucking things up?I was only using i2v from the other model...
>>109502574here's a tip: if you have to fantasize about how you'd be living in a splendorous perfect a future once the bubble bursts, then there is no bubble. this is like saying>i cant wait for the inflation bubble to burst so i can be rich again soon!
I wish H3 had that magical second pass LTX had that mad the output look perfect.
I am what the late /r/ called a "wizard". You are merely a boy.
>>109502591>>109502625Just look at the trajectory of where AI is going. As local AI continues to improve, reliance on cloud services will decrease, making many of those services increasingly difficult to sustain. MiniMax has already made a huge dent in the video genning space, while China has also made a significant impact in the LLM space.From the beginning, China's strategy has been to disrupt the American cloud AI business, and that strategy is becoming more effective as its models improve. The bubble bursting is inevitable, as local AI continues to close the gap with cloud services
>>109502623nta but if the prompt is good and the source clip/positions generally match whatever you're transferring it onto, yeah
>>109502398>https://files.catbox.moe/rda7zi.mp4>someone finally used this audio i linked many threads backSOME GOOD SHIT RIGHT THERE IF I DO SAY SO MYSELF I DO SAY SO
>>109502645It apparently has it, but minimax didn't release it yet. Some kind of upscaler to reach 2k from the advertised 720p of the model.
if the AI bubble will burst (it should, there's no way to sustain its unprofitability) it will sweep the entire world's economy, and we'll have a great reset or something like that
>>109502574give me your copium and hopium dealer's name, I need some too
>>109502402Apparently it's fucked (checked with gpt pro model).
>>109502664Yup, and it's happening in exactly 2 weeks.
>>109502657lol. you realize that you have to pay more than a cloud subscription if you want to run competent AI models locally, right?>but muh stable diffusion and sageattention!nobody cares about image/video, it's a nothingburger. it doesn't matter if deepseek beats gpt or claude, you still need millions of dollars in hardware to run it at full potential which is why local consooomers will never have affordable vram. openai could crash tomorrow and there would still be millions of B-list companies lined up to buy their hardware before you ever get a chance at it.
Pulled latest comfy and spectrum and now there's like a weird desync between the in-comfy progress and sampler preview and the step count outputted in the terminal. Like, in-comfy's displays are lagging behind the terminal steps, and the default sampler preview is all stuttery and fucked, or doesn't play.
>>109502398>>109502408>>109502438OH YEAH here's the video i was talking about that someone did in ltx. full 20 seconds of coherent super high quality video one could even say it's some GOOD SHIT RIGHT THERE MMHHMMHHMMhttps://civitai.red/images/130079342https://files.catbox.moe/0uptmb.mp4
>>109502679>spectrumdelete this cope shit already
>>109502574The US dollar is probably going to be Zimbabwe-tier in a couple of years. A loaf of bread will be $1k dollars.
>>109502664Every other day someone predicts the AI bubble burst, and it's still going since at least 2023, with more investment than ever in robotics, various models, and especially in memory (from colleagues, there is a gigantic scramble to get new production units running to make this by 2027).
>he pulled
>>109502367But I did?
>>109502367video gen is more interesting than image gen
>>109502574>>109502688Kek you people are delusional.
>>109502690Yeah, they're burning your retirement money.There is no profit. The money will run out
>>109502574You realize we are in the niche, right? Normalfags who use AI pay for SaaS/cloudshit. Local is utterly irrelevant, if it wasn't, companies would never open-source useful models, but since there are so few people capable of running them at acceptable speeds, they don't really lose their margin of profits, it's no big deal and they also give others the opportunity to improve their software stack. I genuinely think the only reason OpenAI, Google, Anthropic and others don't release weights is because they don't want competitors to investigate how they trained their models and got them so good, other than that no one but large datacenters are able to run their models, so I doubt they are concerned that a rare rich autist wouldn't pay for their API because he can run the weights local (I bet almost no one in /lmg/, even the richfags, runs Kimi K3 regularly at home with full precision).And also, if selling retail GPUs somehow cannibalizes the still-profitable datacenter business, Nvidia and the others will not give a fuck about us
>>109502574Do not listen to him. If you don't buy your hardware now, you'll regret it in a few months when things will be even more expensive. No one wants you to own things, not the gov, not openai or nvidia. They want to tie you to a cloud subscription and making things expensive help them achieve that goal.
>sol + cache>res_multi / simple 25steps>5min20secgod I love cope nodes. the quality is impeccable
i remuxed all my files to try and save space and in the process lost all the prompts so i just deleted them which was probably healthier
>>109502574At this point I’m pretty sure this is just what things cost now and your job simply makes them seem expensive.
Someone make a ps2 graphics lora for Anima
>>109502667Yeah I've no idea, just tried it with their Turbo wf (which is just Turbo 2nd pass, first gen is regular) on a different prompt and I'm getting smearing so I'd have to use the normal model to see its benefits
You said that int8-convrot vae is good, but vmaf says there's a difference between fp16 and in8-convrot, and quite a big one (vmaf).
>>109502709This gen is missing Will Smith with all those spaghetti on the floor
>>109502722I tested it a few times myself and the speed savings were barely more than the fp16There was weird smearing and artifacts too so I decided to leave it alone and work on other ways to gen faster
>>109502702Sure anon, sure.
>summary:>[video continuation]>detailed_description:>ahh ahh mistress
>>109502703I'd say we are in the same position as the game consolefags who are complaining that Sony is abandoning physical media and going all-digital. They probably looked at reports, determined that most people buy games digitally, knows that selling digital copies is more profitable than allowing people to use second-hand physical media, and decided they would go all-digital.Running AI on gaming GPUs is similarly a drop in the ocean when compared to cloud usage, so there is no reason to pander to us (yet)
ok. hear me out.>0.4mp first pass>upscale to 0.8mp>4steps turbo lora at low denoise.
>>109502736Well, I disabled sage-attention to test it. Two videos, same seed. There's a, like, 0.2 difference in vmaf between two videos in common. But between int8 vae and fp16... Don't use it anons
https://old.reddit.com/r/StableDiffusion/comments/1vjaq5e/minimax_h3_pinkcherry/
Buying a house right now is retarded. In a year or two, the real estate bubble will burst, house prices will come down and banks will shift their focus back toward the middle class market. We'll probably start seeing big houses with much larger amounts of land. I'm talking like 64 acres for 100k.
whats the veredict on low step loras?
>>109502709Awesome
>>109502760thanking you for contribution sar
>>109502747kek, good times
>>109502763Scheisse
>>109502679I got that problem earlier too, anyway i switched to the turbo lorahttps://files.catbox.moe/kq0nw8.mp4
>>109502763>whats the veredictvideo okay saarbut audio very bad
BLAME is cool but has anyone done any Berserk animations yet?
>>109502768>again, no offense
>>109502780Berserk already his plenty of real animations tho. BLAME! only has that one shitty netflix movie.
>>109502773audio working fine on the ema600 checkpoint and no shift node either
>>109502788Nothing past Golden Age has good animation though.
>>109502771>vram debug nodewon't it load models from the disk each time the generation starts? my nvme is fast, but still raping disk isn't good
>>109502788>adapting the entire thing in the 97s styleholy kinolifr though, it's fucking depressing that such a beloved anime got a bullfuck adaptation, meanwhile absolute slop gets a budget
>>109502810Was testing out if it solved the hitching issue at the start of gens. It did actually decrease the overall time
H3's days are numbered>>>/wsg/6210574
https://www.reddit.com/r/StableDiffusion/comments/1vj3l5b/minimax_h3_spectrum_v021_new_offline_replay/updated
>>109502800well yeah, notice I didn't say anything about the quality of those animations lol.There's a bunch of reasons I'm going with BLAME!Thematically it's already pretty closely related to AI generation. The art already lends itself really well to black&white (color matching would be hell). And Nihei's pretty loose with the designs of the characters. lots of differences from page to page which makes the character inconsistencies between gens kind of mirror that.
>>109502402>>109502716Ah, featherweight stack workflow https://matlowai.github.io/ComfyUI-MAINodes/#featherweight, trying it right now
>>109502821Not reading that sloppa, if I wanted to read claude I'd just ask it
Finally my GOONINATOR 3000 can have video slop
>>109502821Good job, it broke the gen preview in the sampler
>>109502832oh also, barely any dialogue.
>>109502816The movies were decent, especially the 3rd. Shame they never got the chance to continue them.
>>109502821qrd?I fucking hate people who unironically paste that kind of sloppa
>>109502856Ask your AI waifu to summarize it.
git: pulleddrivers: updatedits kino time
>>109502821>Yeah, I'll fix that in the next patch, just wanted to get the release out of the door asapSo you push it out with one of the most critical features, a fucking preview of the video, not working. Are you simple?
>>109502835>Speed: on our card the w4a8 pipeline ran somewhat slower per clip than int8 (29 vs low-20s minutes for the 5 s case); on a 32 GB card that is the wrong comparison, because int8 does not fit and offload-thrash costs far more.Nvm, going to a different quantization than int8 is just dumb
>>109502870There is still hopehttps://files.catbox.moe/p00k43.mp4So far I've been using just sage. Maybe the time for improved generation can be cut in half with cache? https://files.catbox.moe/p00k43.mp4
>>109502687>delete this cope shit alreadySpectrum turns a 13 minute gen into a 7 minute one for me, and the results are pretty much visually identical
>>109502863workflow: BROKEN
>>109502712Tempting. Would be even moreso if you provided a dataset.
>>109502866it's fixed as its at 22 now
Anyone tested if nvfp4 text encoder really breaks coherency and overall way worse than int8?
>>109502927i'm on the latest and its still happening for me
Where are the dance videos?
I need like 20 5090s to get all the ideas out of my head.
non_diegetic_music: N/Ayep, it's quiet time
So does Sigma Shift + Spectrum fuck up handheld camera movement? I think it does for me. Like sometimes it will produce a vibrating camera.
>>109502933Haven't compared but it MAY make a difference for the ref2va model. In my tests, it is struggling with large prompts
>>109502883idk how people prefer cache over spectrum just look at thje two side by side on the right
>>109502941bro I had to fire up the 5090, the 6000 alone wasn't enough. It's unreal the reference shit.
>>109502863Isn't there someone you forgot to ask?
>>109502953Why is she squirming like a retard
>>109502942>integrated_multimodal_description: N/A>overall_soundscape: N/A>non_diegetic_music: N/A
>>109502953post new gens.
>>109502952speed niggabesides you're probably gonna need to regen anyway so speed is always better
>>109502954It really is. It can be tricky to figure out exactly how to prompt for the thing you're trying to render, but once you get it figured out... holy fuck...
>>109502954>I had to take the Porsche out for a spin, the Ferrari just wasn't enoughmeanwhile I am happy using my Ford Fiesta 1.4
>>109502952Biggest difference i see is a forward roll vs a side roll and she still gets ripped into two bodies in both
>>109502402Motion is not always better, but it does get rid of the smearingBaselinehttps://files.catbox.moe/o4r3jh.mp4De-ropehttps://files.catbox.moe/grs5xa.mp4Slo-mohttps://files.catbox.moe/crn09s.mp4Flux 3 prompt I was trying to mimic (though in its case I used a simple prompt https://files.catbox.moe/rf653a.mp4), even its API has this smearing issue that the MIT paper fixes.
>>109502965>>109502942not for me bitches.https://d.uguu.se/mlzEHovp.webm
>>109502987Actually, motion may be more funky because I took extra steps on the de-rope side
I still think wan2.2 is overall better than h3. The main thing h3 has going for it is sound.
>wan2gp added frame injectionok, ok, now we are getting somewhere
>>109503001Patrick don't you have to be stupid somewhere else?
>>109503001You're crazy. H3 understands physics and object interactions that I never managed to get Wan to do.
>>109502730>straight out of diffusion.https://files.catbox.moe/sddmtd.mp3>Look in thy glass>Shakespeare's 3rd sonnet.>>109502742>ace step 1.5 xl base, in case it's not obvious.>This is the style prompt:>jazz. angry screaming female singer.>This is how I structured the lyrics prompt. idk, "virtual singer" doesn't seem to do much, but idk, I'm constantly throwing in random seasonings to see if anything is interesting. idk there's bass at the end, so maybe that part worked.
>>109502993yesterdays gens reheated. :vomit:
sota local voice cloning and music gen when?
>>109503008>>109503010go and try wan fun-VACE then come back.
Local bros, we can do local on the cloud now>https://videocardz.com/newz/modders-gain-full-windows-desktop-access-on-geforce-now>free tier cant be exploited>but paid versions of geforce now can>hijack it and install lm studio>5080h (RTX 5080) at their disposal
>>109502978oh anon, both the porsche and ferrari get GAPPED by a 4 door family sedan. It's time to realize EVChuds WON
>>109503024i'm never trying anything wan ever again you fucking retard.
>>109502937Here yoyu arehttps://files.catbox.moe/44kp8a.mp4
>>109503022Literally the only things we're missingI had great luck with omnivoice though
>>109503048Can cloud models even do sota voices yet?
>>109503001>>109503024I have Wan fatigue. Every video model since Wan has used its architecture in some form. I'm never touching that shit again. H3 broke the cycle and is now my hero
>>109502993ZAMN! SHE'S 14?!also please post your other ones, i missed them and noticed the reposts in the collage.
>>109503022We already have SOTA music gen (with LoRAs), and since Alibaba went closed source and likely won't release Qwen Music, we now wait for ACEStep v2 to answer that question for a music gen base model.
>>109503061>ban evading
>>109502574It's probably more of a 'damned if you do, damned if you don't' situation. I don't think that there will be a simple bubble burst, but more like a monkey paw one: a new GPU technology that is needed for a new type of AI, and while we can then probably buy the old GPUs for cheap, the new cool thing will be pricewalled behind the even newer and even more expensive new advanced GPU types and the ones who bought now will also be unable to use the new technology.
>>109502367Minimax has no gray safety box failures, and despite you setting up whole scenes, it's easier to prompt than Ideogram
>>109503027>local on the cloud
>>109503027cringe and poorpilled
>>109503065Other whats? Dancing Mays? I've got about 45, it's all basically a long troubleshooting process to figure out the best way to generate rhythmically synchronized (sexy) dancing.
check this out, reference can copy the style of old ps1 games and add a character in the style.https://files.catbox.moe/do0ecs.mp4
thats ok, ill get bored of ai eventually and go back to playing video games instead so i dont need to worry about buying a super computer
>>109503061lmao besides the context of the vid yeah seedance 2.5 does look good
>>109503092yeah, those cool dancing may figures. share whatever you want. i just thought they were hot enough to hotglue tbqh.
Seedance is rumored to be a 200b param model, so it HAS to look good lol
you can use turbo lora as an upscale pass.
>>109503061Okay it looks good lmao
>>109503097yeah, it's kinda grim btw. nta, but a couple of days ago I got a hold of a seedance. And h3 is way worse than seedance in terms of video quality, camera movements, and... welol quite everything.but h3 is local, and that's great.
>>109503061bruh.>seedance cloud shitFucking rogue employees
>>109503027>>109503087>not wanting to make dumb shit on custom to geforce now hardwarengmi
>>109503061>1 hour to gen>8 hours gooning like the pedo he is.
>>109503137You mean 8 hours bypassing the safety filter. I sure wouldn't want to get caught with that
the end cutscene where she makes the "yuck" face and sound is Oscar-tier>>109503113>Seedance is rumored to be a 200b param modellook what they need to do to mimic 1.5x our power>>109503120>Fucking rogue employeesactually like 60k of stolen keys iircthat vid probably cost at least $1000 on API to make>>109503119>h3 is way worse than seedancebut the thing is, i don't NEED seedance 2.5 to make a girl posing, or basic gooner shit. so there is actually a "good enough" threshold for a significant chunk of my uses with AI and I think H3 crosses that, and there's no going back from that>>109503137kek, but seedance clips at 720p take around 3-4 minutes so i believe it. idk what the rate limits were fior him since the website he used thats watermarked in the corner switched to whitelist-only so maybe he was doing 1 video at a time, maybe 4 videos at a time idk
>>109502604I keep hearing this and I keep trying to insert a single mom I know into deviant shit and can't get it to work to save my life.
is he right or is he right?https://files.catbox.moe/ie449b.mp4
>>109503103Alright here's a randomly curated selection I guess. You'll be sick of Toxic like I am after watching them lol.https://n.uguu.se/luFBwOUd.webmhttps://d.uguu.se/wXcBLqqH.webmhttps://h.uguu.se/GuryEQBp.webmhttps://n.uguu.se/vGKpqGwC.webmhttps://n.uguu.se/oqghfnYa.webmhttps://d.uguu.se/DcTnvsxi.webmhttps://h.uguu.se/WPRWMazd.webmhttps://d.uguu.se/hBCKiFGA.webmhttps://n.uguu.se/AKPveQhW.webmhttps://d.uguu.se/RhgwbbBN.webmhttps://h.uguu.se/qieBhMBN.webmhttps://n.uguu.se/iKnSLvly.webm
>>109503061>>109503097>>109503113>we're gonna have shit like that locally in 2 yearsfuture's looking bright, maybe less so for my wallet with how expensive upgrades are gonna be, but trust the plan.also lol'd hard when she saw his dih that was actually pretty well made despite the content (which i FULLY disavow for any glowies reading this)
>>109502220>rtx 6000 propost workflow please
>>109503168nothing shitdance can do, can compare to what the reference minimax model can do. you can do LITERALLY anything, which is something a rigid, censored, static model like shitdance can't do.
>>109503027is this some sort of poorfag joke i don't understand?
>>109503168>I think H3 crosses thatfair enough by the way. the only things I make is gooner shit. And it's better to use local models for that. Hope that 2k upscaler they have will make wonders.
>>109503169>insert official ref model prompting guide into claude>generate skill .md file for a prompt enhancer>feed skill to gemma 4 ablit>tell it exactly what you want it to do, and exactly what inputs you have to work with>check prompt afterwards and adjusting timing/content to taste
>>109503061>genning this off localAbsolute balls. I'm scared even using grok to fix my prompts.
>>109503193Can I use skills in llama-cpp web UI?
>>109503174kino, thanks.
>>109503200A skill is just a system prompt. Use whatever. I use LM Studio myself.
For anyone using ComfyUI, when I divide a line into two, is there a way to use switch to decide which line I want to go?
its a bumpy ride
>>109503174fuck off spammer
might switch to int8 vae, VAE encode is FUCKING SLOW
>>109503174Why are you obsessed with pokemon girls?
autism alert:>>109503215>>109503211
Sulphur is at $9855/$10,000. That was fast.
>>109503119Yes because you are a Vramlet shizo that use mini max with 0.4 turbo slop lora and eassy cache. In 2mp the model can easily reach that level of detail and you know, the model can do the part you corpo cannot.
>>109503213but it makes outputs look like shit. chill man, it doesn't even make it faster.
>>109503224never underestimate coomers
Is the Pokemon autist and the Rocketgirl fag the same person?I was hoping he finally realized no one outside of /vp/ cares about that shit at this point
>>109503197>Absolute balls. I'm scared even using grok to fix my prompts.it was with stolen keys. do not try this at home.>I'm scared even using grok to fix my prompts.look into zero data retention openrouter endpoints or ask AI to explain them to you. these are what companies use for provable compliance for processing healthcare records and stuff so its actually private. if you trust openrouter's providers that say they don't keep logs at all you can just use openrouter and any model you want and pay with crypto
>>109503225you forgot that minimax h3 also has an api version
where is kino?
>>109503229It makes the decoding faster, but yeah, it also lowers the quality
>>109503232Or just run things locally so you don't have to trust a single third party
>>109503241>where is kino?im out of ideas and getting tired of first person pov for sfw stuff>>109503251i was gonna suggest that too like qwen 3.6 or something but you're probably running h3 locally already
>>109503256>im out of ideas and getting tired of first person pov for sfw stuffmind sharing the prompt? i wanna start doing foot soldier kinos
>>109503225>In 2mp the model can easily reach that level of detail and you knowHow much VRAM would H3 need for like a 15 sec video at 2mp?
I'm jealous of you anons, I'm so autistic I'm updating everything properly, and checking every custom node for security issues before even trying h3.
>>109503268yes
>>109502604I used grok to gen every member of blackpink but with fat asses then superimposed an image of myself in them and have been going hog wild with i2v. Seriously feels like you're a god. The shit I've made them do is diabolical.
>>109503270>checking every custom node for security issueswhy do you need to do that? what are you up to my man?
This is 100% what seedance 2 doeshttps://matlowai.github.io/ComfyUI-MAINodes/#featherweightFixes the blurryness completely
>>109503270Anything to report yet?
>>109503270i use someone elses ui so all i only have a single git history that i review before i update
>>109502333why can i not have a asuka or hermoine gf can someone pls explain why does the universe not allow me this?
>>109503279im gonna cry
>>109503279someone translate this Claude speak into actual human words
>>109503241is in /sdg/
With R2V using an audio reference for voice, do any of you have the issue where an undesired noise/vocalization plays at the start of the video?
debo this is no time for jokes
>>109503297you wish
Had no idea 4chan now takes mp4
>>109503266>mind sharing the prompt?ask your AI for live action gopro stuff e.g.Live-action, cinematic, first-person POV, a wide GoPro-style lens frames a windswept riverside meadow where armored barons in chainmail stand in a grim ring around a campaign tent, banners snapping overhead. >>109503270gotta offload your autism to an AI
>>109503270>>109503280If someone put some secret backdoor into a node that phones home I think it would have blown up already. I fully expect glowies have some undetectable shit already embedded into every possible base local install.
>>109503113It doesn't even look that much better though, the shills will just be shillshttps://xcancel.com/magic_ai_skill/status/2084437127618826257#m
>>109503313thx
>>109503268more than 40
>>109503321And that's the thing. Twitter etc... are filled with Seedance shills, you can tell right away because they do not prompt engineer the model they're comparing against and claim it automatically loses. Flux 3 is better than Seedance 2.5 and is far less slopped, there's no doubt about it.
which comfy "get image size" node actually shows the numbers in the node, so you dont have to connect more shit to it?
>>109503313also if you need ideas, do first person view of big disasters like volcanos erupting or something like the hindenburg crashing down
>>109503335Yep, no contest, F3 is the video model and it's likely magnitude times smaller than Seedance https://xcancel.com/isocialwebseo/status/2085820290437697803#m
>>109503350>video modelbest* video model
>>109503350flux 3 dragon looks like how to train your dragon dragon
>>109501448that's how I became a rebetikochadhttps://www.youtube.com/watch?v=WpDZ1uVDbt4
People seem to be completely forgetting these:1 - People are running H3 local with shit the severely lowers the quality of the outputs, 0.4MP, those "speed optimization" nodes, quants etc2 - They haven't released the 2K upscaler yet3 - That movie-like quality Seedance has may be achievable with a competent fine-tune. No has has done a large scale fine-tune on H3 yet. Training on real high-quality movies may add the "grainy" texture, make things sharper and reduce the "plastic" look. H3 out of the box is pretty slopped, but the underlying model is good.
>>109503384this exact same post was made about wan 2.2 when it first came out
>>109503384Oh, and training on movies can improve the facial expressions too
>>109503384I make high quality works see>>109503308I only gen at 1mp and upscale with rtx it's a cope until they release the upscale but we have flexibility. The turbo loras are also fine, we're like week 1 and look at all the crazy shit we have
>>109503402Wan2.2 was and still is really good. The problem is that the poorfaggotry made it look worse than it actually was.If Seedance 2.5 got open-sourced, most people's gens would look far worse than the API as well
>>109502604How good does this work exactly? I tried it a bit but it just kind of cut to a slightly altered version of the same video. Even with generated prompts. Can you make so it just copies the motion loosely and apply it to an entirely new location/angle/character?
>>109503418did you define the motion as its own subject? then you say its attribute_transferred to another subject
>>109503412>If Seedance 2.5 got open-sourced, most people's gens would look far worse than the API as wellit wouldnt because poorfags wouldn't bother with it because H3 exists
this is legit amazing. I have not tried continuing video beforehttps://www.reddit.com/r/StableDiffusion/comments/1viwvuq/better_avoid_saul_2_minimax_h3/
>seedancehttps://files.catbox.moe/1mo32o.mp4
>>109503430okay i still want you to kill yourself for continuing to link reddit, but credit were its due this is fantastic.
>>109503442saaar we all use the reddit here and we LOVE it
>>109503436ah yes the classic youtube gamer, "the barely perturbed gaming geek".
>>109503215they're very fuckable
>>109503224My only issue is if it will look good or if every donor added the same "professional porn with botoxed milf" shit.
Which of the speed-up options for h3 won't kill my gen quality? I'm just running sage attention 2.2 currently.
>>109503458I gave a bunch of real amature stuff
>>109503430I'm guessing those were genned at maximum resolution with no sage or any type of speed cope cause they're pretty good
>>109503275Yeah, my personal files, secret in env and so on.>>109503280So far so good.>>109503313>gotta offload your autism to an AII use gpt pro model, pretty amazing model, but I verify everything it writes, and it takes time.
>>109503461i think none of them are bad on their own, but if you stack all of them up then it starts to get worse
>>109503461
Haven't installed H3 yet. Do weights actually NEED to be in vram, or can I offload to DDR5 for longer gens?
uhhhh...animator status?this is just with the turbo lora toohttps://files.catbox.moe/uhlb1k.mp4
>>109503308crazy gen
>>109503465Nice.
>>109503461Use the pruned model, and outside of sage 2, literally nothing.
>>109503474>voice actors still safe
>>109503308man, h3 is magiclooks indistinguishable from some old animeAI porn slop spam is about to have a huge quality increase
>>109503472Should I not be running the H3 Mem Eff node?
>>109503492well, voice actors that don't have degenerative brain diseases anyway.
>>109503503depends on your vram but the normal one is better objectively
my brain feels funny ever since I started genning with minimax.
>>109503514I thought he was just Welsh
>>109502367You can't even draw a piece of bread with ideogram with a simple prompt. If doing something even so mundane and simple requires esoteric bbox fuckery, it's not worth the effort.
>>109503519weird innitbeing able to translate both your brain's thoughts and dickvision into a 1:1 video anyone else can see
Had to wage slave and took 2 days break. So what's the meta for quality/speed balance? For me its KJ Sage patch. All other cope caches seems to just make faces and hands blurry during dancing.
I want to use the turbo lora as an upscale pass, but fuck VAE encode is like 2min at .8mp
>>109503471>>109503472>>109503490Thanks anons.
>>109503534If you're on a 50 seriesQuality degradation, especially on higher mp, is almost non-existent
https://litter.catbox.moe/bxk84q.webm
>>109503541Thanks, I'll load up some gens and report back.
>>109503541>sol AND spectrumpost results, i find it hard to believe stacking the two doesn't result in degradation.
https://www.reddit.com/r/StableDiffusion/comments/1vj79sz/minimax_h3_some_test_spectrum_lightx2v_larryvrh/optimization comparisons. It looks like larryvrh Ema v4 600 LoRA is the best over all speed vs quality. Spectrum is slightly better quality but is slower
>>109503551Sure, let me gen something sfw first
We also can tell if you're frauding anon, me and >>109503551 will know
>>109503541i still have no idea if sage attention 2 works on amd or not
>>109503554>heretic text encodersFUCKING DIE
so h3 is what people who can visualize things in their head feel all the time? literally imagine a scenario and see it like some kind of magical performancedamn I'm glad I was born to see what it looked like
>>109503060>sotawhat's that?
>>109503535>apply film grain>no need to cope with uprez/detail fixing
>>109503546I am inversely disappointed by pokemon games vs pokegirls' designs. They're always spot on.
>>1095035192 minutes for 5seconds feels too long.
>>109503554Using ablit TE reduces the quality of your outputs. It also doesn't "uncensor" them. All it does is give you worse results than you would have had with the proper TE.
>>109503558i await with (master)bated breath.>>109503575wow he's literally me
>>109503554Is that plug and play if not I'm sticking to the one KJ put up.
>>109503571You are missing separating the layers into r, g, b and adding subtle noise to each layer and then combining them back together
>>109503546FUCK OFF FUCK OFF
>>109503570state of the art
>>109503586https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/blob/main/minimax_h3_turbo_v4_step600_ema.safetensorsIts just this. Its way better than the light lora / spectrum
>>109503569Yes but my visualization is like an old tube TV. Now we can generate things in high def.
Does Spectrum have an effect on sound quality? Specifically dialogue.
>>109503571turbo lora rapes coherence and prompt following.this way I can gen at low res with zero cope nodes to get maximum model performance. than upscale to get rid of all the ugly artifacts.
>>109503597I'm sure there are people who can imagine in 4k or something
>>109503594idk this shit doesn't work for me at all. is it just bad for ref2video or something? Can't get it going
>>109503601not this one>>109503594
FRESH>>109503609>>109503609>>109503609>>109503609
>>109503599Used to for the ref model, but the new version fixed it or heavily mitigated it. Sounds good now. Seems it broke the preview though, on both the sampler and KJ's custom preview for Minimax. Least it's not working for me and a few other people looks like, least on ref. Haven't tested base yet
>>109503575I don't think that's what "eating ass" means.
One sec lads, I'll bake the non-debo thread.
>>109503608Yes this one >>109502449
>>109503606it needs this loaderhttps://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
>>109503611>trollbakeagain?
>>109503594I don't want to install a custom node, or is he just providing one for people that have no idea what to do?>>109503611autistic miserable wretch
>>109503624mrbones.mp4
>>109503615make a fucking collage. if you can't I'll make one.
>>109503626you need it atm, pruning loras does not work well
>>109503629You have 2 minutes, autist.
>>109503611get a job sharty, stop wasting your time with that inane shit
>>109503631I'll stick to the kj one, I don't really mess with nodes not made by a few trusted sources
>>109503637ok its coming
>>109503640you usually come that fast?
>>109503643no, but the guy he's blowing does
inpainting ace step 1.5 xl base. Loads of fun.Doing Sonnet II of Shakespeare.I reckon I'll do them all, if I live long enough to finish it out.
>>109503637
>>109503646lol lmao
im learning. what do you think of my first gen for h3?https://files.catbox.moe/rdfbyu.webm
stfu faggots you do not disrespect bakers unless he adds his own gens to it
WHOEVER IS VIBECODING SPECTRUM, FIX THE FUCKING BROKEN PREVIEW
>>109503639https://litter.catbox.moe/j6ge0eohu0r678hj.mp4
its up>>109503609>>109503609
>>109503651way past cool bro
>>109503210based vanilla coomer
>>109503649was #8 posted here? holy fuc
>>109503648reminder this week's gensindia diss:https://files.catbox.moe/dioqb9.mp3sonnet IIIhttps://files.catbox.moe/sddmtd.mp3sonnet Ihttps://files.catbox.moe/w2ysvt.mp3
>>109503210>>109503663nvm found itgodtier
>>109503659gotta go fast
Move>>109503671>>109503671>>109503671
>>109503649>give a priest the job of driving one naillmao accurate
>>109503660Remember when all it took was a nice pair of boobs?
>>109503676Yeah, if a single lady with huge boobs and of modest face says hi to me at church tomorrow, I'll punch her. Disgusting whore!
>>109503676just be a teen again
>>109503656>when you look into the workflow and there is more spaghetti than ten ai-generated Will Smiths could eat
>>109503541spectrum is supposed to be at the end of the chain, dummyhttps://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3#workflow-placement
>>109503907Thanks, I'll try it
>>109503430>this is legit amazing
>>109504222