Discussion and Development of Local Image, Video, and Audio ModelsPrevious: >>109948307https://rentry.org/ldg-lazy-getting-started-guide>UIComfyUI: https://github.com/comfyanonymous/ComfyUISwarmUI: https://github.com/mcmonkeyprojects/SwarmUISDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineageWan2GP: https://github.com/deepbeepmeep/Wan2GPNeural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel>Checkpoints, LoRAs, & Upscalershttps://huggingface.co/modelshttps://huggingbay.xyzhttps://civitai.comhttps://civitaiarchive.comhttps://openmodeldb.info>Tuninghttps://github.com/spacepxl/demystifying-sd-finetuninghttps://github.com/ostris/ai-toolkithttps://github.com/Nerogar/OneTrainerhttps://github.com/tdrussell/diffusion-pipehttps://github.com/kohya-ss/sd-scriptshttps://github.com/kohya-ss/musubi-tuner>Qwen Image 2.1https://huggingface.co/Qwen/Qwen-Image-2.1>Minimax H3https://huggingface.co/Comfy-Org/MiniMax-H3https://neta.art/use-cases/en/h3-1000-prompt-list>Animahttps://huggingface.co/circlestone-labs/Animahttps://animastyles.thetacursed.comhttps://tagexplorer.github.io/https://animadex.net>Krea 2https://huggingface.co/krea/Krea-2-Rawhttps://huggingface.co/krea/Krea-2-Turbohttps://lumenastrum.github.io/clio-style-preview/gallery/>Kleinhttps://huggingface.co/collections/black-forest-labs/flux2>MiscShare Metadata: https://catbox.moe | https://litterbox.catbox.moe/Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusionArchive: https://rentry.org/sdg-linkCollage: https://rentry.org/neo_collage>Neighbors>>>/aco/csdg>>>/b/degen>>>/gif/vdg>>>/d/ddg>>>/e/edg>>>/h/hdg>>>/trash/slop>>>/vt/vtai>>>/u/udg>Local Text>>>/g/lmg>Maintain Thread Qualityhttps://rentry.org/debohttps://rentry.org/animanon
Blessed thread of frenship
>>109951187did u use turbo for that video why is it slow motion like
No getting hyped over "coming soon" posts below this line. We will wait until the weights are available to download before making any judgements like respectable good boy genners ________________________________________________________________
guess newfrens dont remember the leadup to h3 release
>mfw Resource news09/30/2026>Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generationhttps://cjeen.github.io/RMD>LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillationhttps://jsxzs.github.io/LIFT>LongLive-Plug: Once-for-All Distillation for Video Generationhttps://github.com/NVlabs/LongLive>MUGEN: Interactive Panoramic World Exploration via Camera Controlhttps://alaya-lab.github.io/MUGEN>LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generationhttps://github.com/PolyU-VCLab/LDMisAE>Honeycomb: Constant-Size Scene Memory Representation for Video World Modelshttps://jackswl.github.io/honeycomb>VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generationhttps://github.com/shim0114/VIF-Bench>RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidancehttps://github.com/yiming-l21/RA-CFGCache.git>LVMT: Video Mask Transformer for Long-term Video Segmentationhttps://www.tue-mps.org/lvmt>CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transferhttps://github.com/0606zt/CLeaR>Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generationhttps://huggingface.co/datasets/Vicky0720/VidScribe>Look Closer: Patch-wise Supervision for AI-Generated Image Detectionhttps://github.com/LF-Jade/look-closer>NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Traininghttps://github.com/JoeZhao527/Noise-Rater>Formalizing the Sampling Design Space of Diffusion-Based Generative Models via Adaptive Solvers and Wasserstein-Bounded Timestepshttps://github.com/aiimaginglab/sdm09/29/2026>MageTrail - V0.3 Update: Modified MageFlow 2.8B Danbooru/E621 Finetunehttps://huggingface.co/RicemanT/MageTrail/tree/main/V0.3
>mfw Research news09/30/2026>Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoEhttps://yuci-gpt.github.io/SplitMoE>Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Modelshttps://arxiv.org/abs/2609.37537>Motion Concept Unlearning in Video Diffusion Modelshttps://arxiv.org/abs/2609.36832>From Scores to Samples: Elastic Forcing for Autoregressive Video Generationhttps://arxiv.org/abs/2609.35491>PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latentshttps://arxiv.org/abs/2609.36199>Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generationhttps://arxiv.org/abs/2609.37407>Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformershttps://arxiv.org/abs/2609.36348>FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generationhttps://fracgen.github.io>ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routinghttps://arxiv.org/abs/2609.37831>Parameterized Stripe Attention for Efficient Video Generationhttps://arxiv.org/abs/2609.37001>Waypoint-1.5: A Real-Time Video World Model for Consumer Hardwarehttps://arxiv.org/abs/2609.37107>Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RLhttps://arxiv.org/abs/2609.37200>RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolutionhttps://arxiv.org/abs/2609.37850>Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy Historyhttps://anonymous.4open.science/w/self-aligned-forcing>Learning via Self-Consistency for Diffusion-based Video Reasoninghttps://arxiv.org/abs/2609.36826>Texture Space Material Diffusionhttps://arxiv.org/abs/2609.37654>SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Videohttps://arxiv.org/abs/2609.37969
>>109951366>>109951379fuck off nigbo
>mfw MORE Research news>Adversarial Training for Pixel Diffusionhttps://arxiv.org/abs/2609.38170>Improved Distributional Diffusion Modelshttps://arxiv.org/abs/2609.37147>GleanVID: Complementary Token Selection for Efficient Video Large Language Modelshttps://arxiv.org/abs/2609.37042>NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generationhttps://arxiv.org/abs/2609.36756>Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Imageshttps://arxiv.org/abs/2609.36882>On the spectral properties of generative denoiser Jacobianshttps://arxiv.org/abs/2609.36210>Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generationhttps://arxiv.org/abs/2609.36364>Reimagine Video Dynamicshttps://yuyuanspace.com/RVD>PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillationhttps://arxiv.org/abs/2609.36638>Reprogramming Vision-Language Models via Structured Prompt Reparameterizationhttps://arxiv.org/abs/2609.36680>HiRAE: Hierarchical Representation Autoencoding with Residual Budgetshttps://arxiv.org/abs/2609.37775>DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losseshttps://arxiv.org/abs/2609.38156>Think Before You Score: Thinking Reward Model for Visual Generationhttps://arxiv.org/abs/2609.37372>What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generationhttps://arxiv.org/abs/2609.37317>BeatDance: Generating Beat-Consistent 3D Dance w/ Hierarchical Spatial-Temporal Modelinghttps://arxiv.org/abs/2609.37400>FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generatorshttps://arxiv.org/abs/2609.37851>Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editinghttps://arxiv.org/abs/2609.36755>Temporal-Attention Head Specialization During Video Diffusion Traininghttps://arxiv.org/abs/2609.31654
>>109951366>>109951379>>109951389
No good qwen training params yet?
any good tutorials on prompts? am trying to generate goons
>>109951422>1boy, large breasts, wide hips, female penis, female insertion, female on femaleyou need more?
>>109951430alr hold on i'll generate this real quick
>>109951430it generated shitty dragonball slop fuck you
cozy breas
>>109951470joe biden would be more of a purple saber guy
>>109951472Say no more
for me, its ideogram 4.5 (unless the prompting is retarded like the previous one)
>>109951535if i cant run it on my laptop its shit
>>1099511876
"Turn this into a realistic photograph taken in Japan"Ideogram gen is uncanny, but it did the job kek.>>109951408>gigantic ifIt makes perfect sense that this model is better than GPT for edits though, you can clearly see it from their preview, how GPT doesn't preserve the quality after multiple edits.
>>109951401Some egg head will figure it out
>>109951551same prompt but in qwen 2.1
Does anyone have a better way of creating PC-98 pixel art? My current workflow is to take a regular illustration and doing post processing to reduce its resolution, limiting its colors, and transforming it to pixel art. This has given me better results than just genning pixel art and fixing it to become actual pixel art. Pic related is an example that I was able to make.
*taps sign* >>109951264
>>109951563
>look at how good the api is anon the local release will be just as good Hm.... where have I heard this before...
>>109951478what are those nipples about
>>109951612Are you the same anon who asked that same question recently?
>>109951643he's flexing really fast
>>109951657Yup. So far the best tool out there has been the pixydust quantizer custom node for comfyui. I also tried unfake.js and spritefusion pixel snapper.
>>109951729
>>109951729https://github.com/jenissimo/unfake.js/https://github.com/sousakujikken/ComfyUI-PixydustQuantizerhttps://github.com/Hugo-Dz/spritefusion-pixel-snapper
>>109951563lol, I just capped a bunch of gigant for source for h3. the only bad news is that it interprets 'giantess' as bigbigonda body style.
Kinoplexatorium status?
>>109951563Looks like my free GPT credits ran out, I can compare with Qwen Image. Turned it into bit of an effort prompt in attempt to make it a bit more realistic, since it seems the model is specialized to do edits in photos/manga/anime instead of cross domain edits like this. Maybe giving it a reference realistic photo might help.>Transform this 2D manga panel into a realistic, cinematic photograph. Maintain the exact framing, pose, and chiaroscuro lighting, showing a young Viking warrior with messy dirty-blond hair and a heavily shadowed face with intense, piercing eyes. Replace the illustrated elements with real-world textures: gritty skin, coarse woven wool, and weathered leather laces. Remove all manga line art, screentones, and speech bubbles, set against a dark, moody background.
https://files.catbox.moe/qyuho9.mp4The plushie version of the Chillet minimax - H3 ref sheet for those that want it: https://files.catbox.moe/61640k.png