[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


[Advertise on 4chan]


File: MiniMax_H3_01252_final.mp4 (3.61 MB, 800x1056)
3.61 MB
3.61 MB MP4
Discussion and Development of Local Image, Video, and Audio Models

Previous: >>109948307

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP
Neural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Qwen Image 2.1
https://huggingface.co/Qwen/Qwen-Image-2.1

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3
https://neta.art/use-cases/en/h3-1000-prompt-list

>Anima
https://huggingface.co/circlestone-labs/Anima
https://animastyles.thetacursed.com
https://tagexplorer.github.io/
https://animadex.net

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/neo_collage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
Blessed thread of frenship
>>
>>109951187
did u use turbo for that video why is it slow motion like
>>
No getting hyped over "coming soon" posts below this line. We will wait until the weights are available to download before making any judgements like respectable good boy genners

________________________________________________________________
>>
guess newfrens dont remember the leadup to h3 release
>>
>mfw Resource news

09/30/2026

>Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
https://cjeen.github.io/RMD

>LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
https://jsxzs.github.io/LIFT

>LongLive-Plug: Once-for-All Distillation for Video Generation
https://github.com/NVlabs/LongLive

>MUGEN: Interactive Panoramic World Exploration via Camera Control
https://alaya-lab.github.io/MUGEN

>LDM-is-AE: Latent Diffusion Model is an Auto-Encoder for End-to-End Image Generation
https://github.com/PolyU-VCLab/LDMisAE

>Honeycomb: Constant-Size Scene Memory Representation for Video World Models
https://jackswl.github.io/honeycomb

>VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation
https://github.com/shim0114/VIF-Bench

>RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance
https://github.com/yiming-l21/RA-CFGCache.git

>LVMT: Video Mask Transformer for Long-term Video Segmentation
https://www.tue-mps.org/lvmt

>CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer
https://github.com/0606zt/CLeaR

>Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation
https://huggingface.co/datasets/Vicky0720/VidScribe

>Look Closer: Patch-wise Supervision for AI-Generated Image Detection
https://github.com/LF-Jade/look-closer

>NoiseRater: Meta-Learned Noise Valuation for Diffusion Model Training
https://github.com/JoeZhao527/Noise-Rater

>Formalizing the Sampling Design Space of Diffusion-Based Generative Models via Adaptive Solvers and Wasserstein-Bounded Timesteps
https://github.com/aiimaginglab/sdm

09/29/2026

>MageTrail - V0.3 Update: Modified MageFlow 2.8B Danbooru/E621 Finetune
https://huggingface.co/RicemanT/MageTrail/tree/main/V0.3
>>
>mfw Research news

09/30/2026

>Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
https://yuci-gpt.github.io/SplitMoE

>Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models
https://arxiv.org/abs/2609.37537

>Motion Concept Unlearning in Video Diffusion Models
https://arxiv.org/abs/2609.36832

>From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
https://arxiv.org/abs/2609.35491

>PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
https://arxiv.org/abs/2609.36199

>Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation
https://arxiv.org/abs/2609.37407

>Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers
https://arxiv.org/abs/2609.36348

>FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation
https://fracgen.github.io

>ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing
https://arxiv.org/abs/2609.37831

>Parameterized Stripe Attention for Efficient Video Generation
https://arxiv.org/abs/2609.37001

>Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware
https://arxiv.org/abs/2609.37107

>Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL
https://arxiv.org/abs/2609.37200

>RelayVSR: Large-Small Model Collaboration for Efficient Real-World Video Super-Resolution
https://arxiv.org/abs/2609.37850

>Self-Aligned Forcing: Streaming Video Diffusion with Differentiable Noisy History
https://anonymous.4open.science/w/self-aligned-forcing

>Learning via Self-Consistency for Diffusion-based Video Reasoning
https://arxiv.org/abs/2609.36826

>Texture Space Material Diffusion
https://arxiv.org/abs/2609.37654

>SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video
https://arxiv.org/abs/2609.37969
>>
>>109951366
>>109951379
fuck off nigbo
>>
>mfw MORE Research news

>Adversarial Training for Pixel Diffusion
https://arxiv.org/abs/2609.38170

>Improved Distributional Diffusion Models
https://arxiv.org/abs/2609.37147

>GleanVID: Complementary Token Selection for Efficient Video Large Language Models
https://arxiv.org/abs/2609.37042

>NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation
https://arxiv.org/abs/2609.36756

>Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images
https://arxiv.org/abs/2609.36882

>On the spectral properties of generative denoiser Jacobians
https://arxiv.org/abs/2609.36210

>Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation
https://arxiv.org/abs/2609.36364

>Reimagine Video Dynamics
https://yuyuanspace.com/RVD

>PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation
https://arxiv.org/abs/2609.36638

>Reprogramming Vision-Language Models via Structured Prompt Reparameterization
https://arxiv.org/abs/2609.36680

>HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
https://arxiv.org/abs/2609.37775

>DMA$^2$: Pixel-space Distribution Matching with Adversarial and Anchor Losses
https://arxiv.org/abs/2609.38156

>Think Before You Score: Thinking Reward Model for Visual Generation
https://arxiv.org/abs/2609.37372

>What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation
https://arxiv.org/abs/2609.37317

>BeatDance: Generating Beat-Consistent 3D Dance w/ Hierarchical Spatial-Temporal Modeling
https://arxiv.org/abs/2609.37400

>FlowMap-OPD: Rollout--Kernel Separation for On-Policy Distillation of Few-Step Flow-Map Generators
https://arxiv.org/abs/2609.37851

>Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing
https://arxiv.org/abs/2609.36755

>Temporal-Attention Head Specialization During Video Diffusion Training
https://arxiv.org/abs/2609.31654
>>
File: output_150_89mb.mp4 (3.69 MB, 1552x2048)
3.69 MB
3.69 MB MP4
>>109951366
>>109951379
>>109951389
>>
No good qwen training params yet?
>>
any good tutorials on prompts? am trying to generate goons
>>
>>109951422
>1boy, large breasts, wide hips, female penis, female insertion, female on female
you need more?
>>
>>109951430
alr hold on i'll generate this real quick
>>
>>109951430
it generated shitty dragonball slop fuck you
>>
cozy breas
>>
File: ComfyUI_01306_.png (1.71 MB, 1280x1280)
1.71 MB PNG
>>
>>109951470
joe biden would be more of a purple saber guy
>>
File: KREA2_00002_.png (1.76 MB, 1280x1280)
1.76 MB PNG
>>109951472
Say no more
>>
for me, its ideogram 4.5 (unless the prompting is retarded like the previous one)
>>
>>109951535
if i cant run it on my laptop its shit
>>
>>109951187
6
>>
File: jackfight.jpg (532 KB, 1792x1792)
532 KB JPG
>>
File: 2121215211101051.jpg (3.71 MB, 4769x3349)
3.71 MB JPG
"Turn this into a realistic photograph taken in Japan"
Ideogram gen is uncanny, but it did the job kek.

>>109951408
>gigantic if
It makes perfect sense that this model is better than GPT for edits though, you can clearly see it from their preview, how GPT doesn't preserve the quality after multiple edits.
>>
>>109951401
Some egg head will figure it out
>>
File: image_00000.png (789 KB, 1152x896)
789 KB PNG
>>109951551
same prompt but in qwen 2.1
>>
Does anyone have a better way of creating PC-98 pixel art? My current workflow is to take a regular illustration and doing post processing to reduce its resolution, limiting its colors, and transforming it to pixel art. This has given me better results than just genning pixel art and fixing it to become actual pixel art. Pic related is an example that I was able to make.
>>
*taps sign* >>109951264
>>
File: QWEN2-1-IMGEDIT_00002_.png (838 KB, 448x1088)
838 KB PNG
>>109951563
>>
>look at how good the api is anon the local release will be just as good
Hm.... where have I heard this before...
>>
>>109951478
what are those nipples about
>>
>>109951612
Are you the same anon who asked that same question recently?
>>
File: KREA2_00001_.png (1.58 MB, 1280x1280)
1.58 MB PNG
>>109951643
he's flexing really fast
>>
>>109951657
Yup. So far the best tool out there has been the pixydust quantizer custom node for comfyui. I also tried unfake.js and spritefusion pixel snapper.
>>
>>109951729
>>
>>109951729
https://github.com/jenissimo/unfake.js/
https://github.com/sousakujikken/ComfyUI-PixydustQuantizer
https://github.com/Hugo-Dz/spritefusion-pixel-snapper
>>
>>109951563
lol, I just capped a bunch of gigant for source for h3. the only bad news is that it interprets 'giantess' as bigbigonda body style.
>>
Kinoplexatorium status?
>>
File: 54454515454.jpg (3.8 MB, 5001x2721)
3.8 MB JPG
>>109951563
Looks like my free GPT credits ran out, I can compare with Qwen Image. Turned it into bit of an effort prompt in attempt to make it a bit more realistic, since it seems the model is specialized to do edits in photos/manga/anime instead of cross domain edits like this. Maybe giving it a reference realistic photo might help.

>Transform this 2D manga panel into a realistic, cinematic photograph. Maintain the exact framing, pose, and chiaroscuro lighting, showing a young Viking warrior with messy dirty-blond hair and a heavily shadowed face with intense, piercing eyes. Replace the illustrated elements with real-world textures: gritty skin, coarse woven wool, and weathered leather laces. Remove all manga line art, screentones, and speech bubbles, set against a dark, moody background.
>>
File: 1762407425508409.png (2.22 MB, 1632x1216)
2.22 MB PNG
>>
File: file.mp4 (2.23 MB, 864x480)
2.23 MB
2.23 MB MP4
https://files.catbox.moe/qyuho9.mp4
The plushie version of the Chillet minimax - H3 ref sheet for those that want it: https://files.catbox.moe/61640k.png



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.