Discussion and Development of Local Image, Video, and Music ModelsPrevious: >>109722945https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GP>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Animahttps://huggingface.co/circlestone-labs/Animahttps://tagexplorer.github.io/https://animadex.net>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/ldgcollage_v2>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
Blessed thread of frenship
>>109730780Tanks 4 bake
>memorial day weekend>plenty of weed >plenty of snacks >plenty of vram >plenty of ideas fellas... it feels so good
>>109730817>memorial day weekendbro living in the past
>>109730829the exact names of holidays have become meaningless to me since i now goon to ai 24/7/365
I got a 5070ti but didn't have the right pcie power cords so it's coming tomorrow.I have +4gb vram now so can I make videos comfortably? Pic related is what I want, I don't need SOTA minimax stuff if that's more demanding.
>>109730780>mfw Resource news09/04/2026>lightx2v/Minimax-h3-Turbo · FL2V Turbo 4-step v1.2 (768p)https://huggingface.co/lightx2v/Minimax-h3-Turbo/discussions/52#6a9a890895a616c64799324f>ComfyUI NVIDIA DLSS 5 Visual Enhancerhttps://github.com/Konohamaru04/ComfyUI-NVIDIA-DLSS-Frame-Interpolation>Viggle-Animate: Character Replacement in Video from a Single Repainted Frame https://huggingface.co/Viggle/Viggle-Animate>DSAQuant: Denoising-Stage-Aligned Quantization-Aware Training for Video Generationhttps://robbyant-research.github.io/DSAQuant>Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuationhttps://github.com/AMAP-ML/StateAgent>FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlowhttps://byeongjun-park.github.io/FlashRender>LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipeshttps://huggingface.co/inclusionAI/LLaDA-Image>ComfyUI-VDN-H3: v1.4.0 — Faster streaming, VRAM-aware buffer retention Latesthttps://github.com/Saganaki22/ComfyUI-VDN-H3/releases/tag/v1.4.0>AetherScale for ComfyUI: GPU-native NVIDIA video enhancementhttps://github.com/vizart-vj/ComfyUI-AetherScale09/03/2026>lightx2v Minimax-h3-Turbo ref2v Lora v1.0https://huggingface.co/lightx2v/Minimax-h3-Turbo/blob/main/minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors>DreamX-Creator 1.0: Model Weightshttps://huggingface.co/GD-ML/DreamX-Creator>ComfyUI MiniMaxH3 CLIPCached: disk cache for MiniMax H3 conditioninghttps://github.com/Mu5hr00moO/ComfyUI-MiniMaxH3-CLIPCached>VDN-Minimax-H3 (VDN-H3): Hybrid Attention to Speed Up Video Models with Near-Lossless Qualityhttps://huggingface.co/OpenVDN/vdn-minimax-h3>SolarWM: Open Data and Scalable Training for Long-Horizon Video World Modelshttps://junchao-cs.github.io/SolarWM-Web>TAME: Temporal-Aware Mixture-of-Experts for Text-Video Retrievalhttps://github.com/sejong-rcv/TAME
>mfw Research news09/04/2026>OctWorld: Long-Range World-Consistent Video Generation with Octree-Based 3D Mappinghttps://maxtirerror.github.io/octworldpage>One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editinghttps://plan-lab.github.io/editvid>Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generationhttps://arxiv.org/abs/2609.03557>ToPO: Token-Conditioned Preference Routing for Attention-Based Latent Diffusion Modelshttps://arxiv.org/abs/2609.03688>SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generationhttps://arxiv.org/abs/2609.03806>Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMshttps://arxiv.org/abs/2609.03820>Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation Systemhttps://arxiv.org/abs/2609.04151>LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipeshttps://arxiv.org/abs/2609.03796>SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolutionhttps://arxiv.org/abs/2609.03813>Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioninghttps://arxiv.org/abs/2609.04183>When Do Frozen VLMs Respond to Image-Free Object-Token Edits? An Answer-Key-Free Protocol and What It Revealshttps://arxiv.org/abs/2609.03429>Who Speaks for the Pruned? Visual Token Pruning as Coverage Optimizationhttps://arxiv.org/abs/2609.03158>The impact of phase information for few-shot fine-grained image classificationhttps://arxiv.org/abs/2609.03829>The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMshttps://arxiv.org/abs/2609.04110>Editable Visual Designhttps://arxiv.org/abs/2609.04034
>>109730896Minimax works on low vram setups like yours but it will be hell. Just wait OR cope with some classic seed variation animations.
How does the first frame in minmax work?if I use different image proportions it tends to make an entirely different composition based on the image, which makes sensebut sometimes even when the image proportions match the video output, it doesn't use the reference image as a first frame
>>109730896You cant wait 1 day? XD need to goon that bad anon?
>>109730924>>109730936
>>109730896You could easily run Minimax with that GPU, but if such simple gens are all you want, you might just use LTX-2.5. It's a lot faster and will probably also give you what you want.>>109730943While it's not a huge amount, I think 16GB VRAM is plenty for doing simple Minimax videos
>>109730998Are you sure you're starting the prompt with the proper syntax per the official skill? You should be giving that markdown file to an agent to have it prompt for you desu.
cozy
When is the new version of GemmaPrompt coming out? no updates for 3 weeks
>>109731007I'm cranking out 40 second one shots on the 5070ti, it's plenty at least at 0.3 mp res. https://files.catbox.moe/k93b34.mp4
0.4*
never prompt photo background, horrifying
>>109731002Seems like the mental health is having a field day. It's weekend after all.
>>109731201aww shiet did this nigga lost some words? i am going to bully this guy with my Kentucky dialect!
>>109731137Psychologists could write volumes about this video
>>109731216>>109731201>upset
fucking h3 loras are so bad
>>109731302skill issue, been making pov anal wet pussy fingering squirting loud tongue out orgasm (all at once) videos all night
>>109731313>No featsBored tonight?
>>109731313silly coomer porn stuff works fine with shitty loras, but they crumble as soon as you prompt something more complex like movement or change of positionsIceKiub = loras by these guy suck ass, I've tried both of his and both have gave me shitty resultsslop twerk lora works wonderfully tho (not made by him)
>>109731302creepy shit
why are this and anime two separate threads
>>109731326h3 loras results remind me of LTX gens, the people who are making h3 loras are probably using the same dataset and same captions as LTX, hence the same shitty results
>>109731335Dev schizo wanted try to make a cult following and weebs just took it over and made it theirs.
>>109731319too old and not obese
>>109731319I haven't seen fennec girl in ages>>109731335a schizo was trying to kill /ldg/, so they spammed a bunch of generals. the anime one stuck around
what we needed, another schizo has arrived, since its Friday night, he's probably drunk as always
>>109731112Anyone?
>>109731350>>109731344nono, there was a legitimate case. during wan2.1 release, the threads were devoid of anime. 1girl realism took over which pushed many posters away.
>>109731354No it was ani seething
>>109731335This is the local thread btw no ugly cloud gens plox
>>109731378He seriously paying to make a gen that shit?ROFL
>>109730780>Reading manhwa >Get idea to train character lora on a character>Check Civitai >Several people beat me to it already https://civitai.red/search/models?sortBy=models_v9&query=Seyoung%20NA>Get slightly annoyed at myself for not doing it firstAm I stupid for feeling this way?
Extend your clips with motion context, it's fun!https://files.catbox.moe/kxxd7a.mp4
>>109731393she's quite popular
>>109731399Pretty cool.How many frames did you sacrifice to keep stitching look ok and preserve motion?
>>109731385He even pays to send his prompts directly to the FBI if you can believe it
>>109724544>Was the Pepe in this one added in a 2nd step or with the original prompt?the latter
someone get astra to fix comfyui
FBI handlers don't often come out but I have been tasked to monitor these threads.
>>109731442What's wrong with it?
>>109731313But what about foot insertion?https://files.catbox.moe/emgq98.mp4
>>109731192kino
>>109731378>>109731385Wow what's this sour grape hostilityI guess I'll take the good stuff elsewhere
>>109731448what isn't wrong with poothon bloatshit?
What should I generate?
>>109731561use imagination
>>109731561Geological events, demonstrating how mountains have been formed.
>>109731555So nothing
Really enjoying NovelAI V5~
a chinese created this scene using h3. i’m not sure if he used ref or t2b, but he used local h3. could someone replicate something similar?
>>109731603you sound like a CEO that enshitifies everything for a quick buck
>>109731446My FBI agent is a hoe.https://files.catbox.moe/8mbwas.mp4
>>109731688They are pretty cool guys. Driven by their ranks and ambititions. Only a certain kind desk jockey monitors internet threads.
>>109731302it works well for me. but your image is tricky because of the phone. generally, it works great, audio sync and shit
>>109731736Welcome back, master. I have generated these same ships.
>>109731719thats horrible, low resolution slopmy problem is with the details for example the right hand of the woman on the right is deformed, also the way the woman on the left turns her body is unnatural, looks like LTX tier, dont delude yourself because your coomer brain is telling you that kind of slop is acceptable
>>109731754we shall assemble a fleet
>>109731719https://files.catbox.moe/nkm5lr.mp4
>>109731771Aye Sir!
>>109731765it's a 3-second video, you faggot. the movement can't be perfect. i'm not going to generate more minutes for you and this blue channel. at least it works. now kys. Next
>>109731393Train a better one, the-n.
>>109731796Then don’t post “it’s works great” when it doesn’t, go back to the cave you came from vramlet coomer
>>109731413Either 1.6 (~30 frames) seconds or 3.8 (~90 frames) seconds at the end of each clip depending on how good it was looking. 1.6 seconds was usually enough (though if you’re not careful with the prompt the sound might get weird)Clips are generally 10-13 seconds (minus the carry over amount)
>>109731854suck my d, you idiot who doesn't even know how to use a lora, lmaoooooooooooooooooooooooooooooooooo
I thought summer was over
>>109731869Thanks anon, will definitely try once I'm not lazy enough to create multiple corresponding prompts.
What is the best for make a 1 minute AI video with ref using audio?
I wish someone did hentai moans h3 lora, they're so much erotic than the shitty dirty talk or carrot eating blowjob stuff.
>>109731870lol you made me laugh tho, it’s ok lil bro
>>109731880It’s a real game changer, just remember to use something that actually carries over the latents (I use a modified version of https://github.com/seitanism/ComfyUI-H3-Motion-Context-MultiRef vibecoded into the UI node you see above). The quality of the stitching is night and day
how do you get a workflow from a video? when I drop it into comfy it simply adds it under a load video node
>>109731897why not just use prerecorded moans as audio input?
>>109731681Still nothing
>>109732158you must be a bad CEO if you make nothing
>>109731993I've seen quite a few context continuation nodes, is this the best one?
>>109732149A lora would be more varied and correspond better to whoever voice timbre I want to use.
>>109732063As far as i'm aware, you don't. At least this is how my ComfyUI install works i'm pretty sure workflow. Metadata is only embedded in PNG files, and there's no current implementation to embed that into a video file. I don't even think there's a standardized way to do that unlike there is for PNG files. You could probably create a custom version of the Save Video node but then, any metadata embedded into the video what pretty much only work on YOUR local install of ComfyUI ( and also ensure comfyui even supports dragon drop metadata instruction from videos in the first place).
Save Video
>>109732184but it would sound tinny. real audio is better quality
>>109732175Retard good-for-nothing no coder, you're supposed to explain why it's supposedly sucks ass instead of being a crybaby dumb fuck artists who expect everyone to bow down to its opinions.
I am seated, yes, at the kinoplexitorium.
>>109732187Huh? I can drop videos into comfy and get the workflow every time. It would be extremely annoying if this didn't work.
>>109731302i used few of such + acts. horror show. avoid them at all costs.>>109730312>>109728048>>109728701which model and what lora is this?
>>109731771 very cool model and lora plz?
>>109732345no lora, base krea2turbo
No other flat chest lora on Minimax ? The one i found in civit is shit
>>109732365i got kera few weeks ago after not considering it at all during release, it is good for styles indeed it is, has sd sdxl feel in gensbasic prompt for the material look in your prompts plz? has star wars vibe
>>109732464as with all my prompts, its quite a mess, so its hard to know exactly whats kicking in for the aesthetic. the logline is>astronomy picture of the day photograph of a spaceship this seems to have a very nice quality of tethering to near-contemporary tech. I later have a mess of>film photography aesthetic, photojournalistic realism, documentary photography qualityhard to say whats doing the heavy liftingI'm happy to give a catbox too if thatd be more helpful
>>109732505ty notes taken will testand catbox sample would be much appreciatedthat is kick ass desktop tier gens material
>>109732559https://files.catbox.moe/ypt31r.pnghopefully there can be some helpful stuff in there
>>109732594it werks!tyvm, yes same style, looks greatwildcards in your prompt, look; not maintained anymore i think, original (llm stuff added though, pull oldest one if you want basic wildcards + lora loading from encode node this pack offers):>https://github.com/Tinuva88/Comfy-UmiAIor check this fork of it>https://github.com/I-ShadowStar/Comfy-UmiAI-Indemnitate
>>109731917These are all really cool. Make one with Gondola too. Should fit the style nicely.
>>109732857i like the idea but i dont think base anima knows gondola
Is anyone even interested in music sloppa? I never see it posted here.
note: the 0.1 ref2v lightx2v lora is better at lower res than 1.0, where you need to generate at 0.9-1.0mp or you get more noise.so 0.1 at 4 steps seems ideal.
>>109732991music sloppa? funny you should askhttps://suno.com/s/SHLLmAWcdsl0MDtg
>>109733006for example this is a quick test at 0.3mp with 0.1. it works, now I need high resolution. 4 image references, works fine.https://files.catbox.moe/za0dp9.mp4
>>109733029That's not local...
>>109733036https://files.catbox.moe/f4n6xj.mp4another test (not the 10s vid)
gem alert?https://files.catbox.moe/y7cns1.mp4
>>109733045tough shit
https://files.catbox.moe/0d2w01.mp4>10 feet behind them is gigachad, a 7 foot tall muscular man in a black speedo who is wearing a black tshirt with the design of <Picture 5>, in a full greyscale style with no color.>it workswhat a fun model.
>>109733171kill yourself
>>109733176go ACK
>>109733178boring
getting closer and closer to the desired result.https://files.catbox.moe/h4esp4.mp4
1mp was worth the waithttps://files.catbox.moe/uez98b.mp4
i meant what i said
>>109733171why is gigachad wearing hp?>>109733280cute
>>109733347Welcome back, Master~!
>>109733171>>109733246>>109733288what is the genuine point of these videos. No wonder nobody takes local serious to begin with. At least troons actually make spicy shit worth watching and rewatching.https://files.catbox.moe/45oeci.mp4https://files.catbox.moe/6ian1y.mp4https://www.deviantart.com/shiftingfun/gallery
>>109733383the point is to make them seethe
>>109733407>themAre "them" in the thread with us right now?
>>109733436given /lgbt/ there are statistically at least a few.
>>109733407You are suffering from self-reinforcing delusions.
>>109733407this feels pointless, most troons are on reddit and discord. The pro ai troons are less ideological extremists compared to the anti ai normie zoomer cattle because their programmed default is already pro trans because its the "social acceptable" position. It's the ai haters and luddites you would want to focus your energy on offending and trolling. They are the ones that will likely get open source and chinese models shut down for good.
>>109733485I dont have a goal, it's for fun. troons getting mad on twitter is just a free serotonin hit.
how good is audio reference in H3I used a 15-second audio ref and told Clanker to repeat the same phrase. It produced a good 5 second then the rest was just yapping with the same accent
>>109733458>statisticallyStatistically it's less than 0.1% of the population and there's maybe ten people posting here.
>>109733501use a clip that is like 10s of audio with no musicthen just say <picture 1> is the reference for Guy, with the voice of <audio 1>then when you prompt dialogue it should clone it. worked for my JC Denton test, etc.
>>109733518>then when you prompt dialogue it should clone it. worked for my JC Denton test, etc.You have to actually prompt dialog? I just tell it to repeat the exact same phrase in the audio. I didn't prompt the dialog
>>109733501just remember to code all your spoken lines with <d>text here</d> otherwise it will add random gibberish lines.
>>109733520is this khroma?
>>109733529oh, you want the gen to copy the audio? in my example it's for voice cloning it, then you prompt what they say and it copies the voice.
>>109733534yea like, the audio ref is a person singing (no music). Do I have to prompt the lyric? cuz "make video of this person sing <audio 1> " doesn't seem to work
>>109733529i'm also not sure actually, desu I haven't been able to get the voice cloning to work. I once tried making a clip of Trevor from GTA 5 say something, even provided a voice sample but he kept sounding like Michael. Not sure if the Michael voice is just hard baked into the GTA 5 token.
>>109733549not sure so far ive just tested voice cloning, but if you supply audio you could just do "sings to the music in <audio 1>" maybe?
>>109733549You have to prompt the lyrics and reference the audio just like in the guide
good talkbatmanhttps://youtu.be/DxMJ0I_JrTQhttps://suno.com/s/9DrJWWt0Ufajrjf9
>>109733549I tried this one as well with very poor results. I provided the song a reference image and a voice sample trying to replace the voice in the original song. I had results where it tried to play the entire song in a 15 second clip. Then I manually cut the first 15 seconds of the song, provided it with lyrics but eventually rage quit.Figuring out the prompt is hard. I'm telling it to use the use the audio 1 as the music being played, audio 2 as the voice to replace the one in the original clip, but it keeps on changing the music or only starting to play it at the end of the clip.I guess i need to look at the prompt again but figuring it all out becomes quite daunting.
>>109733533nope its just krea2
>>109733623looks very clean
cozy breas
>>109733647t
>>109733647
>>109733687nice
why isn't she naked?
okay now im satisfied. ive also confirmed the 0.1 ref2v lora (lightx2v) at 8 steps, euler/simple, is ideal. the 1.0 one for me has more noise even with euler. I wanted to test multiple things:>Use <Picture 1> for the physical identity of gigachad.>10 feet behind them is gigachad, a 7 foot tall muscular man in a black speedo who is wearing a black tshirt with the design of <Picture 5>, gigachad is in a full greyscale style with no colorwhat a model. ref model in a sense makes it unnecessary to have loras, which are good but in this case the picture reference is the lora.https://files.catbox.moe/ofa2x0.mp4
>>109733637its a nice model but its a shame the community at large won't give it a chance. This space has become too fragmented between vramchads, 16gb mid-vramlets and poor vramlets. Krea2 lora section on civitai is depressing to look at. Its truly the human centipede of boring and nasty fetish slop with a very few goodies buried in there. Minimax h3's lora section is even worst to look at.
>>109733716what do you mean? krea 2 is the standard for image gen, klein edit 9b for edits, and minimax for video. krea 2 is a great model.
>>109733716idk man to me the krea2 loras look pretty good. Lots of cartoon styles, which is the kind of stuff I like. What exactly are you looking for because to me it seems like the Civitai Krea area is bussin' (no cap)
hm...
>the camera hard cuts to <Picture 1> showing troon diving at high speed towards the water as the camera tracks them, speed lines show them traveling extremely fastyou can use an image reference as a keyframe. kinohttps://files.catbox.moe/4ofydr.mp4
This gen was fun https://files.catbox.moe/77k3u0.mp4
>>109733821better phone screenshot:https://files.catbox.moe/jkyzys.mp4
>>109733737i want more character loras and not just styles and fetish nsfw stuff flood the page on civitai. illustrious page on civitai has some much good character loras to play with unlike krea2. It just agonizing and painful to see how pale the krea2 lora section is. Come on anon, who the fuck wants to download and use a lora of a fake non-real photorealistic ai influencer model OC of another user. May be my expectations are just too high and I'm out of touch with the kinda stuff people like in the space. I know people are hungry for character loras judging by the numbers of downloads for several of the ones I've made and uploaded so far on civitai.
What stuff with AI can benefit from having 2 (or more) graphics cards? LLMs can load models across GPUs to benefit from higher VRAM; can image gen or video do the same?
>>109733926Minimax supports tensor parallelism and so do many other models as well. Though you probably won't come far if you have to ask
>>109733926you can gen 2 1girl at the same time
>>109733926looking for that one lucky seed out of five hundred that is barely usable in H3 gets faster if you have more GPUs