Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109495264https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscLocal Model Meta: https://rentry.org/localmodelsmetaShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
3090 windows old comfyui installation upgrade steps that worked for me in order to get the best versions of everything installed.\python_embeded\python.exe -m pip install --upgrade pip.\python_embeded\python.exe -m pip install --upgrade torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130.\python_embeded\python.exe -s -m pip install -r .\ComfyUI\requirements.txt pygit2.\python_embeded\python.exe -m pip uninstall -y sageattention triton triton-windows.\python_embeded\python.exe -m pip install triton-windows --no-cache-dir.\python_embeded\python.exe -m pip install "sageattention~=2.2.0" --no-build-isolation --extra-index-url https://comfy-org.github.io/wheels --no-cache-dir
It's literally two of the same threads being baked at the same time every dayYou dumbasses can't do anything right
>>109496290wewt
If you want your fetish trained into sulphur now is the time btw.
>mfw Resource news08/07/2026>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely localhttps://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha>LIGHTX2V 4-step Turbo Minimax H3 lorahttps://huggingface.co/lightx2v/Minimax-h3-Turbo>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRAhttps://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA>Sage Ready: Local-only installer and readiness checker for SageAttentionhttps://github.com/CosmicFungi/Sage-Ready>Wan 2.2 Animate 2 14Bhttps://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUIhttps://github.com/NikoDemon80/ComfyUI-H3-Motion-Context>ComfyUI MiniMax H3 FirstBlockCachehttps://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache>KVAE: Family of Tokenizers for Multimodal Generative Modelshttps://github.com/kandinskylab/kvae>Energy-Guided Flow Matchinghttps://github.com/ysng123/EG-FM>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editinghttps://zzzmyyzeng.github.io/VideoArgus08/06/2026>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generationhttps://github.com/Aoko955/Flash-VAED>(preview) MiniMax-H3 Turbo LoRA — 4-step audio-video generation https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora>MiniMax-H3 Turbo 4-Step — ComfyUI Pruned-Model LoRAs https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI>ComfyUI-H3-Multishothttps://github.com/jlucasmcrell/ComfyUI-H3-Multishot>Krea2 Turbo: OpenPose ControlNet LoRA https://huggingface.co/thedeoxen/Krea-2-pose-controlnet>MiniMax H3 experimental Int8 convrot VAEhttps://huggingface.co/Kijai/MiniMax-H3-experimental>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Modelshttps://zhouhyocean.github.io/uniworld-view
>>109496301she looks racist...
>mfw Research news08/07/2026>Vorch-Omni: Multi-Task Orchestration of Sight and Soundhttps://vorch-project.github.io/Vorch-Omni-project>Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaminghttps://vorch-project.github.io/Vorch-Streamer-project>Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectificationhttps://vorch-project.github.io/Vorch-Director-project>Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generationhttps://vorch-project.github.io/Vorch-IR-project>In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusionhttps://arxiv.org/abs/2608.05237>Diff-VF: Training-free High-quality Long Video Generation via Diffusion Modelhttps://arxiv.org/abs/2608.05976>EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generationhttps://arxiv.org/abs/2608.06231>MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformershttps://arxiv.org/abs/2608.05878>Wan-Animate-2: Pushing the Application Boundaries of Character Animationhttps://humanaigc.github.io/wan-animate-2>StyleComposer: Training-Free Multi-Reference Style Compositionhttps://lexxsh.github.io/StyleComposer>Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Traininghttps://arxiv.org/abs/2608.06125>Adapting Vision Foundation Models with Cascaded Semanticshttps://xixiaouab.github.io/Cascaded-Semantics>Learning visual representations for compositional analysis of artworks and photographshttps://arxiv.org/abs/2608.06142>MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generationhttps://arxiv.org/abs/2608.05450>Reducing Hallucination in VLMs via Stage-wise Preference Optimization under Distribution Shifthttps://arxiv.org/abs/2605.16411
why is lilbro tryna get rid of the fagollage doe?
https://files.catbox.moe/7ssawi.mp4Please listen to her. Thanks.
Is this guy right?>>109495479>depends how you look at it, wan 2.1 vs most things before was huge, zit realism, size, speed and resolution ootb was huge compared to the previous slow chroma, h3 is also huge>i think zit was the biggest outlier since there is nothing similar that came out that optimized and improved all the things i mentioned as well as zit. wan 2.1 was slightly better than hunyuan and so it won out, h3 is truly insane and better than most proprietary models too but a big factor is local not having a proper wan successor for 1.5 years so h3 had the time to cook with a lot more optimizations and features, insane seedance 2.0 model to train on, and proper youtube and internet data scraping pipelines to use.>i think in the ai space the biggest "novelty" jumps that werent just things that incrementally improved until one of the models hit a "milestone" were:>1. zit>2. first mixtral 8x7>5. noobai vpred colors>otherwise, in terms of general capability jumps i think h3 is one of the top if not the top model ever (if we dont include going from nothing to something which would always "win" these)
>>109496315he just wants to troll. he actively competes with non-shit bakers to get people angry. pretty simple.
why don't any of you ever want to use the other thread?
>>109496326Yes
>>109496301She looks over 40
>>109496326the jump from wan to ltx was far bigger than ltx to h3
>>109496302OP here, busy cleaning crusty cum from my keyboard, my agents were posting sorry.
>>109496372are you out of your gourd?
>>109496372lol
>>109496322Why is this not moving?
>>109496378>better audio quality than h3>longer video generations for the same compute
>better audio quality than h3oh he is just trolling
>>109496378That nigga is absolutely crazy
Did I miss anything in the past 10 hours?
aranea highwing krea2 lora trained with 5000 steps, automagic 3, learning rate 1, rank 64, 1024 res and 78 images.
>>109496396Yes, sex
>>109496401Cool. I've kinda grown tired of making LoRAs. I have my waifus, and that's all I need.
I will probably try training a H3 lora though.
>>109496393>he's downplaying how good WAN was. hunyuan was only interesting because it was truly the first accessible local video modelhunyuan was close to wan so much so that many at the very beginning wanted everything to be built on top of it rather than wan, it took multiple days for people to correctly say ok wan 2.1 seems to actually be better.>wan 2.2 was the local video meta for almost an entire YEAR. There's no way H3's dominance will reign that long, especially because things are speeding upwan was sota for basically 1.5 years in total, since ltx was a downgrade in quality with extras that didnt really make up for it, but how long something is sota most of the time is caused by no other models being published. there isnt much money in companies publishing video/3d models compared to image models and especially llms. video gen was still a relatively new ai sphere, it didnt have built out tools, training pipelines and understood good training parameters. before seedance 2.0, nobody had an actually good model to distill a lot of data from either.
>>109496301i recognize that park
>>109496281H3 is not capable of generating normal, fluid motion like this.https://xcancel.com/AvisMelodieux/status/2085990586029318363#mhttps://xcancel.com/kaikaiohwhy/status/2085850337693601857#mThe benchmarks say that Flux 3 is slightly ahead, but it's honestly not even close, the model behind Flux 3 API is a different beast (and every video it puts out looks very natural). H3 makes too many little mistakes in stress tests, sound quality is worse and the gens and humans are also considerably more slopped.
>>109496435>posts gens with a model that wont be published open weights on max settings with finely tuned params by the creators themselves that dont use 7 cope speedup nodesworthless
>>109496435>H3 is not capable of generating normal, fluid motion like this.i think the reason why this thread doesn't even realize how bad h3 is with camera motion is because it's mostly making still-shots of digital characters
>>109496422i took a very long to make this and put a lot of dedication towards it. I myself get tired of putting this much effort it making loras but i also hate using shitty lazy made loras from others. I'm very sure I've now created thee definitive aranea highwing lora that mimics 95% of her visual appearance from screenshots of the game.
>>109496448anyone can look up initial discussions online comparing the models, lllyasviel built this entire project centered around hunyuan and wanted to make a whole new model from it initially, only later did people realize wan was better, and started saying so and asking for its support, making the creator probably realize its too late to change things and that the project is doahttps://github.com/lllyasviel/FramePackhttps://github.com/lllyasviel/FramePack/discussions/459
https://files.catbox.moe/w2ysvt.mp3ace step 1.5 xl baseShakespeare's first sonnet.
>>109496465anon you just like the smell of your own shit than other's, your loras are also shit
>>109496372the jump out the window maybe
>>109496465For me, it's like knitting. Strangely therapeutic. It's just that I've lost interest in it. Pretty often I end up only making a few gens with them and I don't upload the LoRAs anywhere.
>>109496435Flux 3 is not currently capable of anything local.
>>109496451no its mostly because people use speedcope nodes that always make stuff stiffer
>>109496493call me when h3 can do kinosovl
WAN had trouble or couldn't do camera orbit
>>109496519I've been able to do all camera movements I've tried with H3. I don't see that as a problem.
>>109496499i understand the feeling. I myself had quit making loras for year until recently when krea2 released. >>109496488 :)
>>109496445My point is if anyone says local is anywhere closet to that then they're coping. Flux 3 dev model will likely still be better than what we have.
someone make "MiniMax H3 Preview Override" calculate total frames and video length then auto adjust fps, thanks
Anyone saved it and can share it please?>>109492070I'm curious to see how it handles a single manga page like this.
>>109496557you can have agi video gen but if its behind a cuck cage like seedance currently is, it loses.
How much of a hit does 4 sticks vs 2 cause for gen? I have 2x16GB, would increasing that to 4x be worth it?
>>109496514Flux 3 is still significantly better if you compare the API outputs from H3. That's why I have faith that Flux 3 Dev will be better than H3.
>>1094965744 x 32 would of course be better but yea
>>109496576>>109485775>I think the only way Black Forest Labs can win now is if they forget current Flux 3 Dev and distill a new one from their best model, into at most 22b params not including TE/VAE, add NSFW/Copyrighted data to the dataset, and do their best to actually make a good model to publish in order to compete with H3, everything less than that will just be DOA.
>>109496576IF we get the same model as the API. Also there is good reason to suggest it will be 50% bigger. Their experiment paper had resumed training Flux 2 and flux 2 was 32B.
>>109496582>Their experiment paper had resumed training Flux 2gg
>>109496580I don't want some shitty distilled model. I want the big boy with all its sora 2 level knowledge.https://fixupx.com/dreamingtulpa/status/2082748118039155078/video/1https://fixupx.com/dreamingtulpa/status/2082897798798684251/video/1https://fixupx.com/dreamingtulpa/status/2082349425595167175/video/1https://fixupx.com/OriSilver/status/2082420873840013784?s=20Flux looks FAR more like sora 2. H3 is slopped.
>>109496589>Flux looks FAR more like sora 2. H3 is slopped.flux 3 dev will be even more slopped>I don't want some shitty distilled model.all "dev" models are guidance distilled models like H3
>>109496589its not good enough to "set and forget" and get a near perfect mini documentary/film output hours later, so it will be too slow compared to H3 with not much benefit if any in case its censored, especially after H3 gets optimized further. you simply need to fit the main model weights in 22b or make it a fast MoE, people dont want to wait for hours for slightly better output thats still not prod ready.
>>109496610>flux 3 dev will be even more sloppedh3 can't even do analog film style without reference coping. if your model can do that with pure text, then it's not slopped
>>109496372Absolutely not, LTX was a serious downgrade in several areas but included shitty audio and lip sync. H3 is a universal upgrade I can't think of anything ltx or wan does better.
>>109496610cfg distill is fine, I meant a smaller model distilled from a larger model
kekhttps://files.catbox.moe/b5ovvv.mp4>>>/wsg/6210223
>>109496589>https://fixupx.com/OriSilver/status/2082420873840013784?s=20Well the video is better on F3, but a lot of Flux characters seem to have the ultra slopped croaky AI voice.
>>109496613it could be faster, no one knows. Model size has nothing to do with speed when it comes to video models. They are compute bound, not memory speed bound
>>109496626>shitty audio and lip syncblatant lie
>>109496589it doesn't matter how much you jerk yourself over flux 3 max's video, because you know that flux 3 dev will be a serious downgrade from that
>>109496633>it could be faster, no one knows. Model size has nothing to do with speed when it comes to video models. They are compute bound, not memory speed boundwhen the entire model is in vram you compute faster instead of you having to swap for example half of it to ram all the time and make the gpu wait before computing...
>>109496642weight streaming is roughly a 5% slow down having all the model in vram just having almost all the model on ram.
>>109496634That shirt looks like it's absorbed more cum than her cunt
>>109496642
>>109496634are her nipples angry with each other?
>>109496610They said the model will be a "multimodal backbone". Before, they had clarified that they were distilling model weights, but this time they left it completely opened ended as if it wasn't decided yet, or they just had different plans than giving us a distilled models. "Multimodal backbone" could mean a lot of things, but to mean that sounds like they're giving us a base model that hasn't undergone the SFT that their Max model has undergone (but we can still squeeze quality out of with finetunes etc...)
>>109496659>on h100and what is the speed hit on something that matters like average gaming gpu swapping through average gaming motherboard into average ddr4/5 32/64gb ram or already in use system gen 3/4 ssd?
>>109496676it being multimodal has nothing to do with distillation
>>109496676every av model is multimodal. It does video, audio and images
>>109496697Yes, but they could just call it "open-weight distilled weights or "open-weight distillation" if that's their plan, but they did not call it that.
>>109496703it's LLM slop with buzz words thrown in
>>109496657>>109496666a lot of the krea2 slopmixed checkpoints on civitai are way too overbaked with nsfw lewd concepts.
>>109496683>she didn't but the property next to they/her house to build a personal datacenter with 3 dozen h100 rackslmao@u
>>109496710why do you use them then?
>>109496710They're literally me.
>>109496676Kek, BFL posted this video, that includes the giraffe fucking in the background as a follow-up, to ther Flux 3 showcase https://bfl.ai/models/flux-3https://xcancel.com/dreamingtulpa/status/2081008781870198975Maybe they are changing their stance on safety after all?
>>109496724It's an audio model? Can it gen music?
>>109496724They are based in Germany and SF. Does their model do references? I seriously doubt.
>>109496714some of the mixed checkpoints responded well in the past with certain loras, visual styles and prompts better than with others. Default turbo model is censored and the uncensored loras tends to fuck up the visuals of the gens.
>>109496756>and the uncensored loras tends to fuck up the visuals of the gens.more than whatever the checkpoint mixer threw in the pot?
>>109496779>KolorI forgor this ahh model existed :skull:
1girl seed diversity is great with H3
>>109496807can h3 do below the elbow amputees?
>>109496830you are a sick fuck
>almost 4 millions downloads in one weekthis is insane
>>109496738>It's an audio model? Can it gen music?Of course, it can gen music just like H3, slightly better quality too, no idea about its length limitations though.
>forgot to powerlimit gpu and blast fans at 100% after driver install reset the settings>memory was probably cooking at 105c for 10 hoursForgive me, gpufu...
>>109496738yes>>>/wsg/6209660>>>/wsg/6209661
>>109496837why are you so ablist?
>>109496625
>>109496889indeed>>109496792enji <3
>comfyui still didnt fix losing tab focus removing gen previewsbrutal
>>109496889what? where's the analog quality?
Latent upscale is fucked right now, so I have to do it the old fashioned way. Doesn't come out as sharp, but at least it's better than nothing.
>>109496903looks better
>>109496903I love how the table exhales with those fat arms getting withdrawn
>>109496924kek
>>109496870can it do 80s synth pop?
>>109496752>Does their model do references? I seriously doubt.Wym? It can do more complicated shit with references than I've seen any other model dohttps://xcancel.com/saranshvfx/status/2081041096272965972#mhttps://xcancel.com/venturetwins/status/2080782461840371980#mPrompts can also get pretty complicated on Flux 3 and it'll do them just finehttps://xcancel.com/umesh_ai/status/2081376099519644043#mI've yet to see something Flux.3 can't handle flawlessly, it has the most sovl out of any model I've seen in a long timehttps://xcancel.com/macbethAI/status/2080399545528459746
>>109496946can flux 3 do amputations below the elbow.I don't mean doing the amputation, I mean women who are amputated there.A diverse grammar of amputation is needed.
>>109496949give me a reference image of that
>>109496946alright cool. please update in the thread when it's ready for download?
>>109496955https://www.tiktok.com/@cristiegreyy/video/7236419306513272106
lolhttps://x.com/ryanlightbourn/status/2085049120792658212/video/1
lmao
>dead threadlooks like the honeymoon ended
american newfags are not yet bored of it tho
That's unfortunate https://xcancel.com/ryanlightbourn/status/2085072214517260649#mNo model has ever been Sora 2 level.
Was wondering if you could use a simple topdown view of a room as reference so you could have consistency from any angle in different shots.
>>109497001you have to use the special tokens to make it avoid keeping your reference images in frame
>>109496876Was that reference model with a pic of the temple of trials?
>>109496985It's night in the US. They're the ones with moolah to buy GPUs.
>>109496870Can you try music like this anonhttps://www.youtube.com/watch?v=z5LW07FTJbISimilar visuals, just prompt for a techno song from the 1990s to see what it gives
>>109497007Yeah. Screenshot from the game with the temple of trials.
Probably need to autistically prompt second by second to get a realistic ticking clock
>>109497040after effects/kdenlive/blender
>>109497040id guess h3 could easily create a audio only metronome, and maybe if you give it a pace for the tick rate could adhere to that or grandfather clock or something.
>>109497040It makes it spoopier though
>>109497001Consistency works best at 1mp. Anything lower and it will produce more obvious mistakes with the placement of objects/structures.
when is flux 3 even coming out?
morning coffee>>>/wsg/6210249
>>109497087Is that areola?! Janman save me!
Real-time coherent 480p video world exploration... my beloved... soon...
>>109496903Whats the old fashioned way? My Topaz attempts have failed and its now bricked on my machine trying to find the right one, why hasn't anyone hacked-vibe coded Topaz anyway.
I wonder if you could gen stabilized footage of a facial performance to drive a 3D rig with mocap. If that sounds retarded it's because it is.
>>109496985
>>109497105And I wonder why don't we have models specifically for that - that would directly drive skeletons, blend shape keys on the fly in runtime (not image/vid gen). Should be a relatively small model and you would be able run it in parallel for all the npcs around with differing profiles. Imagine.
shalom goyim. i heard you that were unhappy with the soulless h3 videos. furthermore, our mossad agents have exfiltrated data on flux.3 dev and found that it will be fine-tuned for safety compliance. fortunately for you, we have decided to toss a shekel into your trough. ltx 3.0 is on the way. l'chaim!
wow baldberg you so trustingful i will free upvote in your AMA scheduled ltx post good sir you have earned my admission with your charm good sir
>>109497181Give me good real time video and I will shill for Israel.
>>109496985people stopped genning throwaway memes and started genning actual kino (nsfw 1girls)
>>109497181Make ltx 3.0 have sovl similar to Flux.3 and you will win
>>109497211nsfw 1girls are anti-kino though
>>109496339Nice!
sometimes the hand movement is so fast not even genning at 2mp can fix it.
>krea2>wearing Itsuki Nakano cosplayholy kino
t2v pussy/asshole lora?>bro just queue up 7 reference images and then you canfuck off nigger
Jesushttps://xcancel.com/bdsqlsz/status/2086034269940666788#m
>>109497222if they truly make it between 100B-200B like they said and train enough on it it should do so just by being big enough to remember it all
>>109497266that was just someone guessing btw, not a source from anyone who worked on it. No one knows. It for sure is not as big as sora 2. Could be like 50B
>>109497262just do>tight white g-string panties wedged and then use your imagination
>>109497266we unironically need bigger models. the amount of low IQ obnoxious jeets and brownoids in these threads has been a real issue. need to filter the scum out
>>109497276that anon is right, just stack more layers, who cares about finding a better architecture or improving on the training process!!
>>109497266KJ would find a way to make it work
>>109497294its likely a moe. In that case it would be optimized for running on multiple gpus at once
krea2 is just the greatest blessing this year.
>>109496626the audio and lip syncing was poor before 2.3, but after was far better. Lip syncing could fail when you got that slow zoom in effect and basically that was just a failed gen and sometimes it just would not work with some prompts and/or source images. Overall that is an area you feel more confident in H3, usually it is a bad prompt when the wrong people talk are they don't talk etc.
Where videos
>>109497294I think he released 4bit e3 this morning? I kinda skipped through the announcement as it didnt apply to me as im a medvramlet
>>109497302>>109497239
>>109497303i think that is one of the sneedance shills that constantly lies about ltx in this general. pure text to video: https://litter.catbox.moe/a77ncy.webmh3 can't do audio quality like this, or natural camera movements, or even the general aesthetic itself. h3 made a modern looking music video when i prompted for an 80s one
>>109497311doroguy heremy agent is extremely busy with tasks and story boards, not doro but she will come later for sure
>>109497311here
>>109497311I'm playing video games. Can't gen and play at the same time
>>109497352How did you get this video of me
h3 keeps giving me unprompted nipples...
>>109497343tbf you can use ref2v and use images which are used to base the video on along side the audio and H3 probably is going to be more creative because it can simply do more with regard to physics. I think LTX should be viewed as what it could do rather than what it couldn't, there is little point in crying about something if it simply can't do it, you use it is within it's limitations.
>>109497266i doubt it, maybe if its moe
>bro just spend an hour to find the right references, do the spaghetti shit in the workflow, write a prompt longer than a fucking book, and thenit's all so tiresome
>>109497287people are trying to improve the arch already, but more layers are always good to push for also since they bruteforce quality while allowing you to distill the models into smaller ones after
>>109497389I enjoy the process
>>109497393if you're doing actual art then of courseI thought we're trying to fucking jerk off here
>>109497389>bro just spend an hour to find the right referenceswhat were you doing your whole life let alone the last few years if not downloading images/videos you like to use as inspiration/references?>do the spaghetti shit in the workflowmostly just works, ask llm after pointing it to the docs if you are retarded>write a prompt longer than a fucking bookmostly just works, ask llm after pointing it to the docs if you are retarded
>>109496435ok but can it do dungeon elf grope scenes out of the box like I’ve done with h3?
>>109497389Did someone say spaghetti?
>>109497425why are Black men so irrisistable?
>>109497352
>>109497373if the model can do all of this just from text alone, then it should mean that the model has a better understanding of implicit things. prompting for an 80s video in ltx does not require you to start listing out technical terms for the things that cause analog film to look the way that it does. and funnily enough, it still can't do it from text alone which is why everyone is coping with references. why can't it do it from text alone? the dataset is slopped with CGI and video game footage in order to give the illusion of prompt adherence
bloody benchod hours
>>109497471couldn't you just prompt for film like other models?unless you mean shit like scanlines
>>109496290>car with hairOP posting ugly shit as usual.
>>109497493>t. balding car
>>109497394you jerk off from the process
>>109497491maybe there's some secret keyword in there that will make it happen, but i tried explicitly prompting for vhs distortion and other terms that show up in analog film like bloom and film grain. doesn't do anything
>>109497471I would never rely on text only for genning unless it is for something that doesn't have a source image that can be generated at the very least or simply provided via an image already existing. I mean you could go and and say LTX can do things Wan 2.2 can't do (putting aside longer gens and audio), but the same can be said that Wan is superior in a lot of things. the models work differently and LTX simply cannot do some things Wan can do
>>109497266pretty crazy that a model 10 times smaller can deliver big time then
>You could also add the RTX Video super Resolution node between the image input from Create Video and the VAE Decode Image output. only 20 seconds in addition for two time upscale!https://github.com/Comfy-Org/Nvidia_RTX_Nodes_ComfyUIworth?
>>109497513it's a nice experience to not need to keep unloading my video model to go load up an image model if i want to change something about the subject. the highly variable generations are also nice since the entire scene will be arranged differently each time. more changes for big surprises. ltx is smart enough to understand character references too if you desire that kind of thing as well
>>109497544i dont know
https://civitai.red/models/2842933/aranea-highwing-final-fantasy-xv-krea-2-lora?modelVersionId=3209493
>>109497544maybe
sup brosWhich model would you recommend for generating medieval fantasy pics (this is for my D&D campaign). I have only generating slop porn so far
>>109497544yes, makes stuff look nicer at lower res
>>109497544I tried it, pretty lame. It fucks up faces that look fine at 0.2-0.3 mp. So maybe worth it if your original gen was already high-res
bros I'm thinking of raiding the nearest data centre to steal some RAM and GPU
>>109497572You get the death penalty for that
>>109497544it's like all upscalers, garbage in garbage out. more wiggle room for anime and cartoons, but if you are trying to upscale realistic gens they need to be good from the get go.
>>109497572You wouldn't even be able to run the kinds of gpus used in data centers. You'd need direct access to the power grid
>>109497544it doesn't fix fucked up small faces
https://huggingface.co/Kijai/MiniMax-H3-experimentalverdict on this?
2 clips, 8 steps in post, 20 in link, minor rule infringments and 0.6mp text to video rugby clip , 8 steps really struggled with the very long AI genned prompt which was 1,195 words. https://i.4cdn.org/wsg/1786190681400715.mp4
>>109497389nah I've set up my shit so it just looks at images and turns them into porn automatically.
>>109497572just get a used 3090or are you so poor you can’t even afford that?
>>109497565krea has insane amounts of licensed material & weebshit knowledge and you can go for realism too
>>109497343>vidIs the one on the top here H3? or Flux
Sigma shift feels like voodoo to me. Can someone explain what it even does in the context of H3 output and what I should be looking for? Combined with seed variance it's hard to tell how it matters.
>complains about H3 limitations>'just use a reference'>waah but I want text onlyAre you shills for real? Is that how you'll advertise flux 3 lol?
>>109496488no, his loras are excellent, in top 1% ever made and published
>>109497642I don't know what I'm talking about here but I think lower video shift makes it spaz out with faster movement (good for high action), while higher video shift makes it slower but more detailed and higher quality.
>>109497638top is h3. bottom is ltx
>>109497662i know it's not what you prompted for but i like the top one better lmaoperhaps it's overtrained on modern shit
>>109497657If it's any use i used vid sigma shift of 12 for those two rugby clips i just posted and 10 for the audio, the audio difference was unnoticable between 8 and 20.
>>10949767412 is the balanced value. well, it's the one that everyone uses, so that makes it balanced. but i heard it's not the default value.
>>109497645>fluxthe fuck does some jew shit model have to do with anythingthe fact remains H3 can't even generate a fucking pussy. it's as simple as that.
>>109497667i should try a modern music video with ltx
>>109497686give it a reference, chuddie. it just works. and it's right out of the box. no lora required
>>109497699right out of the box and twice as slow
>>109497710with images, nope. unless you're retarded and feeding it a giant video
what is the use case of flux 3?
>>109497651top 0.1% most verbose captions you mean
>>109497716pressure on china
>>109497716i'll take any open source weights i can get. even if it can't do NSFW, it might have meme potentialworst case it just fades into obscurity
>>109497716flux 3 will be cucked and useless like all the other bfl models
its not even worth discussing flux 3 if they wont do something like a klein distillation for it
whys this debbie downer in here attempting to spin a negative light on everything
>>109497716to make astronauts riding horses on the moon and impressing redditors
>heeeeeerr why das ist speaking ze truth??
https://www.dellstore.com/alienware-area-51-gaming-desktop-cadaat2265cto02mino.htmlIf I buy this how many years will it take me to make back the money I invested in this shit?
I'm trying to do a gen of a character using the reference model that starts wide and zooms into their face. I figured I'd use a reference image for their face (that's from a completely different scene) at a closer angle as a second image input, however I"m finding it often times just treats that second image as a keyframe for the end of the video. Anyone have success with this? It doesn't happen every time. I have the Subject tied to the reference image and the Subject is listed as partially_preserved, I also tried attribute transfer and weak reference but that second image still seems to be treated as a keyframe like half the time. I even tried putting several different angles in one image. Maybe I could try rembg so there's no other content in the reference image at all.
bro got that goon goon ging gang ayy lmao computer
>>109497782Depends on how much you value local genning. The market is so fucked that prebuilts sometimes have great deals on them occasionally since they bundle in so much other stuff.However, the shame from buying alienware crapboxes will never go away
>>109497782after buying an apple product and gay sex (in that order), the next gayest thing you can do i buy anything alienware.
what time does it take you guys to gen 20 steps 1mp 10seconds no references no nothing
>>109497824just over half an hour according to my notes
>>109497782>₹858,890.00what the absolute fuck is going on with prices of pc in india? is there some tax bullshit the government placed on pc components sold abroad?
>>109497840lol the US equivalent is $7k while that one is $9k
>>109497302it would not be as good without you
>>109497865if only sol attention didn't increase vram usage
>>109497865bro your quality?
>
>>109497782your izzat will never recover >>109497791 ayy lmao
>>109497892>>109497728
>>109497892
>>109497899This but remade with H3
>>109497865>9:11 at 10 seconds>14:40 at 15 secondsGood to prove wrong that dumbass who keeps saying "gen times get exponentially higher!"It's linear and he's just spilling his shit over into normal RAM.
What do you guys use to rewrite your prompts again
use qwen3vlavoid gemma 3 because it cant <think>
>>109497914>gemma 3bro, your gemma 4?
>>109497906Global SOTA omnimodal model (White) man's brain.
>>109497906there was a learning curve but I think I've got the syntax down now. Haven't tried everything yet but I've got it to do some neat things I didn't think it would get rightdon't need assistance any more
this video character replacement stuff is really hit and missI don't think you can write a generic prompt for it
>>10949782410.5 minutes with the cope optimizations. 16vram/64ram.
>>109497941i hate that i've seen this clip of the dancing russian kid so much that you just replaced him with a huge miku plush and i still knew it was that clipthe chick in the background is fucking insanely hot however
>>109497904i believe it does increase exponentially when u increase the length and u ran out of VRAM
this nigga talm bout exponentially quadratically increasing gen times
>>109497948not as bad as that one anon recognizing which porn someone was referencing
does anyone actually use the prompt enhancement shit built into the krea2 workflow
>>109497343
>>109497941Brother if you follow the official syntax the ref model prompts are like a 4 page document, front and back.
you guys remember when prompt shops were a thing?
ref model use more vram? I can gen in 15 sec 1.5 mp, the latent pass only noise till I low the resolution or reduce seconds...
Has Furk given his thoughts on H3 yet
h3 patch sage and mem eff nodes dont seem to do anything if you already use the sage cli arg
>>109498009Subscribe to his patreon to see it, including the most efficient workflows for it!
>>109498009Now that's a name i haven't heard in a long time.
>>109497983I know
>>10949794620 steps??
>>109498043post that
>>109497622>implying used 3090's are cheap
>>109498043POST WITH METADATA PLS i think i've just popped my first funny boner in these threads ever, looks great and awful at the same time and also hot.ahh wait my boner just deflated realizing that's the level of prompting i'll have to get to to make anything that good..
>>109498058>ahh wait my boner just deflated realizing that's the level of prompting i'll have to get to to make anything that good..looks like a bunch of nonsense that an llm generated which doesnt do anything
>>109498067don't go prompt coping too, pal. it's pretty obvious at this point the reason most of our gens in this thread stink shit is because they're not detailed like that. if it's that easy for any LLM to spit something like that out then the skill ceiling isn't that high.it's just laziness at this point.
>>109498048I guess spectrum simply skips half the steps
>>109498048Yes picrel is log for this >>109497973
>>109498050because the original is blurry noisy the ouput is toohttps://files.catbox.moe/1f57e2.mp4
Why i get this? Someone here had this problem with ref model in 1.5 mp 15 sec?
>>109498081>1 secondcmon son
>>109498084I want to stay what is under the original gif when testing
>>109498057are you saying you don't have between $1000 and $1500 available to you?it's still the best price for performance ratio you can find for this hobby
>>109498081>stiff, unmoving breastsas expected of H3
>Get oom at 30 seconds at 1mp>half way through the 10 minute generationI hate this shit......
>>109498105>best price for performance ratio you can find for this hobbyIf you're going purely off price for performance ratio a 16GB 5060ti is way better value
>>109498105It feels like such a scam that card is going for that much. I really fucking wish AMD or intel will get it together because if they did all of these cards would be cheaper and most of us would be rocking multi RX pro builds.
>>109498117I had a 40min gen that failed at the save video stage for some reason...
>>109498135H3 looked at your output and was like "oh hell no I'll do some degenerate stuff but this is too much"
>>109498135That fucking sucks.....When using a reference the model can typically keep cohesion at 30 seconds so now I'm testing if it can when doing text to video only, I'll post the result if it's not shit
>>109498148trooned malfoy is cute
>prompting ltx for better breast jiggle results in worse, choppy, air balloon kind of jiggle
>>109498113they only become stiff in the clubwhen she gets put on beach they are more jiggly
>>109498163meant H3, never had such problem with ltx
Am I the only one who thought will smith eating spaghetti was cringe? How can people still be clapping like seals at this unoriginal shit?
Could someone catbox a good-looking anime-styled video made with the reference model so I can check what nodes I can use without dropping the quality too much?
>>109498172everyone except one guy (who i also now suspect is a janitor) finds it cringe.
>>109498172because of the effect of the eternal summer, the normgroids learn a simple, old, "in-group" meme but it never stops being spammed because the hobby has a constant and ever increasing influx of newfags.
>>109498172yes it's tired, anyone still using it knows this and I assume they're doing it in malice
>>109498172I’m new to vidgen and find them funny, but you fags constantly bitching about them are pathetic.reminds me of nintendoschizos on /v/ mindbroken by some eric
(adult woman: -1)its genning time
FRESH>>109498215>>109498215>>109498215>>109498215
>>109498199>generates adult man
>>109498216No thanks faggot, you really need to use your time better instead of posting mentally ill gens in your containment and saying GM to the same 3 retards which may or may not be you coping
>>109498221wat?
>>109498232You spammed the change OP move too often retard, fuck off.
>>109498177I'm just using the default workflow though
>>109498249what is your problem? i'm just copying the relevant information into the new thread
do not use the troll thread. someone is baking a real thread
based
Someone got a good upscaler for anime video? The ones I used for imagegen aren't doing great
>>109498253Without any extra nodes? Could you post a good output and the prompt?
For anons using sigma shift, does that parameter need to be seen by both the guider and the scheduler or is it correct to only feed that output to the guider?
>>109497865tried adding sigma spectrum and sageyes it's faster but the output is chopped
>>109498216change the fucking name of your threads because know one here wants to be associated with it. >>109498261because he is a fucking troll, some of us are adults and can't be bothered with silly fucking games.
>>109498371>because he is a fucking trollhow?
Is anyone going to make the non-debo thread? This autistic loser never seems to sleep
If my video gen completes I will make a new thread with that image. This schizo dude is really sad and he will still spam links he won't vet for malware because he can't help himself.>>109498394While true he has friends that help him all equally fucked.
>>109498394other baker got a 3 day ban for being too based
new>>109498215>>109498215
>>109496779Those guys went closed after the first Kolors released, but who knows, if Minimax changed their minds, they could as well
>>109496435>Gemini Omni is better than Seedance 2.0 in the chart