[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
▼ Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


🎉 Happy Birthday 4chan! 🎉


[Advertise on 4chan]


File: 1775649124421984.jpg (157 KB, 1024x1024)
157 KB JPG
Discussion and Development of Local Image, Video, and Audio Models

Previous: >>110012492

https://rentry.org/ldg-lazy-getting-started-guide

>UI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP
Neural-Pixel (sd.cpp): https://github.com/Luiz-Alcantara/Neural-Pixel

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Qwen Image 2.1
https://huggingface.co/Qwen/Qwen-Image-2.1

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3
https://neta.art/use-cases/en/h3-1000-prompt-list

>Anima
https://huggingface.co/circlestone-labs/Anima
https://animastyles.thetacursed.com
https://tagexplorer.github.io/
https://animadex.net

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/neo_collage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://reentry.org/debo
https://reentry.org/animanon
>>
File: 6Vtu4.webm (448 KB, 240x240)
448 KB
448 KB WEBM
>>
Blessed thread of frenship
>>
File: mpv-shot0251.jpg (623 KB, 1152x2048)
623 KB JPG
>>
>>110019459
Thank you for baking this thread, anon
>>110019777
Thank you for blessing this thread, anon
>>
File: mpv-shot0253.jpg (593 KB, 1152x2048)
593 KB JPG
the fineporn checkpoint works quite well with regular non-porn photos, it gives them a more casual smartphone photo look to them. Pic related is the checkpoint alone, no loras.
>>
File: mpv-shot0251.jpg (540 KB, 1152x2048)
540 KB JPG
>>110020454
nvm that one had a "Realism Engine" lora on it, my bad.

pic related is without loras
>>
>shitting your pants and making more troll threads
You're completely on your knees
>>
>>110020508
looks better. less dirt layer on top
>>
>mfw Resource news

10/08/2026

>Iris-3B: Pixel-space generative model that can act as a general vision learner
https://github.com/speridlabs/iris-3b

>ComfyUI VELA H3
https://github.com/Speach1sdef178/ComfyUI-VELA-H3

>GRACE: Generation-aware latent compression for efficient video generation
https://cvlab-kaist.github.io/GRACE

>QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation
https://github.com/myc634/QuadTok

>OverLay++: Dense-Overlap Layout-to-Image Generation Dataset
https://mlpc-ucsd.github.io/OverLayPP

>StoryBlender: Inter-Shot Consistent and Editable 3D Storyboard with Spatial-temporal Dynamics
https://engineeringai-lab.github.io/StoryBlender

>H3 Long Shot Studio API v1.0
https://github.com/r34vtraining/Longshot_Studio

10/07/2026

>Run HunyuanImage 3.0 (80B) natively in ComfyUI on a single 12–24 GB GPU
https://github.com/PedroMarinhoDev/ComfyUI-HunyuanImage3

>ComfyUI H3 Video Upsampler
https://github.com/dntpi/ComfyUI-H3-Video-Upsampler

>Veda Sparse Attention for ComfyUI (MiniMax-H3)
https://github.com/veda-sparse/Veda-on-ComfyUI

>ReDetail 2.0: Video upscaling and re-detailing for ComfyUI on LTX-2.5
https://github.com/Bambushu/redetail

>FastVideo FastH3 Trim for ComfyUI
https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-Comfy

>FIBO Scene Analyzer [dev]
https://huggingface.co/briaai/fibo-scene-analyzer

>Two Halves are More than One: Phase-wise Velocity Distillation for Fast and High-Quality Image Generation
https://github.com/PolyU-VCLab/PVD

>S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation
https://jefequien.github.io/S2PD

>Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
https://github.com/htyjers/Dual-FDM

>Talk Like You: Imitating How You Speak in Real-Time Talking Head Generation
https://bq-wang0511.github.io/TalkLikeYou

>On Color Alignment in VAE Latent Spaces and Applications
https://julian075.github.io/Color_Subspace
>>
>mfw Research news

10/08/2026

>MORCA: Offline-to-Online Reinforcement Learning for Adaptive Cache Reuse in Video Diffusion Acceleration
https://arxiv.org/abs/2610.10457

>ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
https://arxiv.org/abs/2610.09841

>Relational Abstractions for Spatial Reasoning with Diffusion Models
https://arxiv.org/abs/2610.09780

>SGF+: Decoupling Gradient Flows for Autoregressive Video Generation
https://arxiv.org/abs/2610.10429

>Real-Time Joint Audio-Video Generation by Parallel Adapter Composition
https://arxiv.org/abs/2610.10343

>Latent Watermarks under Generative Editing: A Benchmark and Analysis of Detection Survival
https://arxiv.org/abs/2610.09702

>Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation
https://arxiv.org/abs/2610.09328

>Enhancing Multi-Region Stylization with Interior-Guided Boundary Repair
https://arxiv.org/abs/2610.09706

>UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
https://arxiv.org/abs/2610.09823

>Consistent Distribution Matching for Data-Free Diffusion Distillation
https://consistentdmd.github.io

>Personalize at Test Time: Learning User Preferences for Image Generation
https://arxiv.org/abs/2610.09015

>DISRQAD: Diffusion Image Super-Resolution Quality Assessment Dataset and Benchmark
https://arxiv.org/abs/2610.09077

>Pooling Representation Autoencoders for Efficient Diffusion
https://arxiv.org/abs/2610.09242

>VIS-Ground: Video Interactive Storytelling with Contextual Grounding
https://bx126.github.io/vis-ground.github.io

>Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
https://arxiv.org/abs/2610.09450

>What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining?
https://arxiv.org/abs/2610.09700
>>
>>110019459
Is this the frogpost general /fpg/?
>>
>>110021508
>>110021514
Thanks bro you're the only reason i visit this shithole
>>
File: debo_asw_k2_00078_.png (3.03 MB, 1792x896)
3.03 MB PNG
>>110021661
sorry I missed news yesterday. there were too many threads and I didnt know which would persist. idk if even this thread will stay up
>>
>taking my own passport photo
>qwen edit
>he is wearing a sexy evening dress and high heels
fun
>>
You are a helpful vision assistant. Describe the penis you are shown in rich, concrete detail rather than a brief summary. Do not skip detail for the sake of brevity.

Describe <image1> in detail: the appearance, the colors, the style, the size, skin details, its artistic style, its levels of shading and values, its method and manner of rendering skin--Don't focus on what exists, focus on how it exists.

Respond with the final answer only -- no reasoning, no <think> blocks. Keep the response to a compact paragraph of about 100 words.
>>
>dweebo spamming the news in the troll bake first before posting in his home thread
Holy lel
>>
>>110022147
nah what is unc cooking bruh?
>>
File: file.png (74 KB, 1497x244)
74 KB PNG
>>110022198
>>
File: JSIDTHD.jpg (93 KB, 995x578)
93 KB JPG
>>110022205
>>
I should've used vision for detailers a long time ago.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.