[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models and Software

Previous: >>109388421

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>LTX-2.3
https://huggingface.co/collections/Lightricks/ltx-23

>Wan
https://github.com/Wan-Video/Wan2.2

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
gm saars
>>
>mfw Resource news

07/27/2026

>Spectral Prior for Reducing Exposure Bias in Diffusion Models
https://github.com/SonyResearch/SPA

>Twins: Learn to Predict Unified Representations with Focal Loss
https://github.com/Tencent-Hunyuan/Twins

>TRELLIS.2 INT8 ConvRot for AMD ROCm
https://github.com/DrBearJew/trellis2-convrot-rocm

>Fooocus Advanced: Updated Fooocus with some new and modern changes
https://github.com/micha42-dot/Fooocus_Advanced

>Prompt Manager — ComfyUI Custom Node
https://github.com/Fictiverse/ComfyUI_Prompt_Manager

>Cosmos3-Super-Image2Video-4Step
https://huggingface.co/nvidia/Cosmos3-Super-Image2Video-4Step

07/26/2026

>Prompt Architect Pro: Uses local Ollama LLMs to analyze raw texts and images
https://github.com/lololerigolo60/Prompt-architect

07/25/2026

>Aozora Trainer: Memory-efficient fine-tuning toolkit for SDXL and ANIMA
https://github.com/Hysocs/Aozora_Trainer

>Mage-Flow repackaged model files for ComfyUI
https://huggingface.co/Comfy-Org/Mage-Flow

>LoRA Dataset Studio: one-tab workbench for the whole LoRA lifecycle
https://github.com/perfectgf/lora-dataset-studio#everything-it-does

>Open Letter: Open Weights and American AI Leadership
https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight

07/24/2026

>SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
https://nvlabs.github.io/Sana/Video2

>JoyAI-Echo x LTX-2.3 — echoVid + ltxAud surgical merge
https://huggingface.co/joeygambino/joyai-echo-ltx23-echoVid-ltxAud-surgical

>Open Dreamer: Open-source Dreamer world-model implementation in JAX
https://github.com/next-state/open-dreamer

>NVIDIA Qwen-Image-Flash
https://huggingface.co/nvidia/Qwen-Image-Flash

>Inference-Time Scaling of Diffusion Models via Progressive Seed Pruning
https://www.vision.caltech.edu/psp

>Spectral Transformation for Layer-wise Global Rank Discovery in Federated LoRA for Vision Transformers
https://github.com/DASS-Lab-Group/SpecTraL
>>
>mfw Research news

07/27/2026

>Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
https://wenchao-m.github.io/ClosetheLoop.github.io

>TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution
https://arxiv.org/abs/2607.22231

>InnoText: A Unified Model for Visual Text Generation and Editing
https://arxiv.org/abs/2607.22101

>Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On
https://arxiv.org/abs/2607.21694

>Rethinking Layer-Wise Information Allocation for Vision Foundation Model Adaptation
https://arxiv.org/abs/2607.21973

>Correlation-Aware and Gaussianity-Preserving Robust Latent Angular Watermarking for Diffusion Models
https://arxiv.org/abs/2607.22386

>FAIR: Feature-Augmented Implicit Regularization for AI-generated Fake Image Detection
https://arxiv.org/abs/2607.22087

>Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation
https://arxiv.org/abs/2607.22034

>Gemma 4 Technical Report
https://arxiv.org/abs/2607.02770
>>
This is the one
>>
Blessed thread of frenship
>>
regarding ace step, if you're using STORK 4, you may find pretty massive cfg to work.

I have found cfg=22 to benefit instrumental parts, still on the fence as to what happens to vocals.
>>
I'm having loads of fun with "lm understand".

You can absolutely get a huge boost to your game using this. You have to um

ok first the way I do it is I do a gen with my good settings. Let's call it "burner song A"

like you download yt music to mp3, or use your method of choice.

bring it into acestep

you use the dropdown on the song and pick "lm understand"

it will do its thing and then you will get the prompt, that's what you want. copied to your clipboard.

then use the menu on burner song A and pick edit prompt

then you need to clear out the seed under flow match parameters. just make it -1

then replace the prompt with the one in the clipboard & gen away. (you may have to adjust length if the style has changed a lot).
>>
anon?
>>
>>109390413
The hell are you talking about man
>>
File: 3496894860846.jpg (352 KB, 720x954)
352 KB JPG
>>
>>109390435
>Discussion and Development of Local Image, Video, and Music Models and Software
>music

ace step 1.5 XL Base is an ai music model. You can use Ace Step cpp to gen with it, and it has Vulkan build support.

Not everything is your boss krea.
>>
>>109390447
Maybe give a little context to your info dump next time
>>
cozy breas
>>
>>109390455
sorry.
>>
>>109390455
and it's like a thing that goes back to the old ace step, reeeeally high cfg and lcm would occasionally produce great results, barely, very briefly. I still don't know why. cfg with ace step is <> cfg with image gen.
>>
>>109390464
btw it says notorious big, but it's like... I don't change what the text says in the song title.

for whatever reason lm understand puts the song title in there, which is silly, you're just borrowing the STYLE.
>>
post your gens that make you do a slight smirk and whisper "based..." while nodding your head
>>
File: 6875975768.jpg (464 KB, 680x1024)
464 KB JPG
>>109390488
this one makes does that for me, but only because the prompt for the gen was
>The photograph features a Korean woman with striking, curly white hair that radiates outward in wild, spiky tendrils. Her face is pale and expressionless, with closed eyes and slightly parted lips.
>She wears a dress with long-sleeved zebra-print fabric on the arms, transitioning into a voluminous, white tulle skirt that creates a soft, ethereal texture.
>>
Other thread has schizo energy.
This is the cozy bread.
>>
>>109390488
catbox is down again unfortunately
>>
>>109390525
based...
>>
>>109390488
https://files.catbox.moe/a8kwr4.jpg
https://files.catbox.moe/6eoq36.jpg
>>
File: 1763636220730005.png (870 KB, 1024x1024)
870 KB PNG
>>
>>109390413
Yeah, I've done that sometimes to get prompts, but I've found Gemini Flash is more accurate than the audio scanning feature so I just use OpenRouter. Btw if you want to capture the vibe of a song you can just ask Gemini, no need to even import the song into it (the prompt it will give you likely will come close).

ACEStep Turbo/Base merge can do interesting stuff with bass out of the box
https://vocaroo.com/16ojvXCw53NZ
>>
>>109390580
1 girl on the piano with the cordial

there, already better than 99.99999% of kreap.
>>
File: 23.png (738 KB, 1280x725)
738 KB PNG
>>
File: vid.mp4 (1.22 MB, 960x640)
1.22 MB
1.22 MB MP4
>>
File: 3643.png (8 KB, 550x53)
8 KB PNG
the kinoplex is over capacity
>>
>>109390488
>do a slight smirk and whisper "based..." while nodding your head
I have never done this
>>
Bump
>>
File: momo_00009_.jpg (528 KB, 1648x1648)
528 KB JPG
>>
catbox down?
>>
>not a single catbox shared
What's even the point?
>>
>>109391476
omg is that... le asian goddess!?
>>
File: momo_00040_.jpg (563 KB, 1152x2048)
563 KB JPG
>>109391592
I don't know, is it?
>>
File: 104143CUI_00001_.png (847 KB, 1216x832)
847 KB PNG
>>
do people avoid sharing catboxes because their workflows rely on custom nodes (and sometimes personal loras) that wouldn't work on anyone else's machine?
>>
sharing is retarded
imagine having all work getting done for you when the fun part is getting there
>>
>>109392008
If you are unable to use the default example workflows...
>>109392022
What work? You have never worked in your life if you think that creating AI image is 'work'.
>>
>esl tard
>>
>>109392030
Cry more faggot. English is my third language. You barely know one. At least I don't squat here 24/7 like you do.
>>
do yourself a favor and post the question and my answer into a LLM so it can translate it to pt-BR for you in a way a 5th grader would understand, retarded esl shitskin
>>
Please sers, .NO more fighting Only good Vibes.
>>
>>109392008
>i just want to gen without thinking about it
use the default workflow.
>my workflow is shit and i don't know how to get better results
post your workflow and people will help you sort it out.
>>
File: 00002-520058723.png (1.83 MB, 1920x1152)
1.83 MB PNG
>>
>>109392072
benchod
>>
File: 00021-2164555992.jpg (554 KB, 1728x2880)
554 KB JPG
>>
File: ComfyUI_6066.jpg (605 KB, 1632x2176)
605 KB JPG
>>
File: 1752286461347948.jpg (222 KB, 720x720)
222 KB JPG
>still no news about local flux 3
i guess this flux is a new api kek, as expected
>>
>>109392497
they're busy letting kekstone and the rest of the useful idiots safety test it first
>>
As always, I’m speechless at just how incompetent Western nerds are.
If we had Russians or Chinese here, wouldn’t we already have a platform where you could trade datasets for GPU power?
You’re an anime pleb with 8 GB and 60 datasets on characters? You like Krea2? I’ve got the GPU, you’ve got the dataset. Win-win

But the faggots here think the 8th trainer or the 567th prompt tool is what the community needs.

In 5 fucking years, I’ve only seen a single community project, and that was back at the start of SD 1.5 - and it failed because the global press went crazy over it. After that, absolutely nothing.
>>
>mfw IL has better image quality and generations than Anima
>>
>>109392584
better start learning mandarin, if you're serious that is
>>
>>109392584
civitai can do exactly what you're proposing with the buzz system
>>
>>109392584
>If we had Russians or Chinese here
ok, they're not here but they're there, so why doesn't this exist for them yet either? the last relevant thing china did was noobai, and that was 2 years ago. if they were competent, we'd have a krea 2 finetune in the works
>>
>>109392584
Obviously all the nons are sapping our energy.

>>109392588
>score_9, score_8
Do Anima tards really?
>>
File: 00077-3902304012.png (2.69 MB, 1856x1280)
2.69 MB PNG
>>109392588
i don't see the appeal for anima whatsoever. the model is a janky mess and the shitmixes are far worst than what's available for illustrious on civitai.
>>109392584
what ai models are russians even making or using? do they even have ai labs and devs over there?
>>
>>109392667
Good to see you again!!
>>
Anima is a lot more fun than Krea2 just because of the better gen speed. I'd probably go with Krea2 if I had a better GPU.
I've said it before but it's like 30-50 seconds is decent and then 60+ seconds starts to feel painful and unfun.

I'm trying to maintain a boner here while I browse this thread briefly before getting back to my gens.
>>
Why ani is spite baking?
>>
>>109392789
never left. working a bonquet (blue dragon) lora for krea2. really wish someone would be based enough to do all the rayman nymph fairies character loras for krea2.
>>
>>109392925
he's trying to fud krea just because a popular workflow includes anima, and that really upsets him
>>
>>109392584
>platform where you could trade datasets for GPU power
because good datasets are valuable, and people would probably make slop datasets to trade.
and for those that have a good dataset, what gpu power would they really need? having a mid-high range gaming gpu is enough to train and gen everything at ok speed, and for any larger training renting shit online is pretty cheap as well for 1 training session
>>
>>109393108
this. it's cheaper in the long run to rent a pro 6000 than to buy one if you need a big gpu for some reason
>>
>>109392466
As a ZITGOD I do want to say that Krea does seem to be sota for capturing character likeness in LoRAs at least.

What training tool/config did you use?
>>
>>109390216
>>109390224
thanks!
>>
least obvious samefag
>>
File: ComfyUI_02038_.png (3.22 MB, 1344x1792)
3.22 MB PNG
>>109393134
You're not a ZITChad if you say that, CHUD
>>
>>109392584
We have 4 or 5 Russians here.
>>
File: ComfyUI_02045_.png (3.23 MB, 1344x1792)
3.23 MB PNG
>>109393172
And we're all ZITChads. No Krea 2 fags are Russians.
>>
nice skin noise kek
>>
>>109392667
Can you unhide your youtube videos so we can all laugh and make fun of you again, sasori?
>>
still dont see why 'anon' is trying to compare krea2 and anima when the models serve completely different purposes. doesnt matter if krea2 can do anime style illustrations as long as it doesnt know danbooru
>>
>>109393172
>russians
nah this place is reddit from the amount of safehorny 1girl feet spam
>>
>ignores all other gens
>only focuses on the ones that have feet
closeted footfag confirmed
>>
File: 151246CUI_00001_.png (893 KB, 1216x832)
893 KB PNG
>>
>>109393296
Perhaps he doesn't understand that specialized models will always beat generalist models at the thing they specialize for
>>
File: 432102931827674.png (2.78 MB, 1856x1280)
2.78 MB PNG
>>
>>109393458
You're saying jack of all trades master of none is fake?
>>
Is Krea 2 dead yet? Is ZIT the dominant force in AI today?
>>
File: 00009-3212878766.png (2.48 MB, 1248x1824)
2.48 MB PNG
>>109393289
i don't intend in unhiding any of them especially in the age of ai moderation of retroactively nuking youtube videos. Not sure why you found those old videos of my younger autistic self amusing to begin with. It's a dead channel like many other dead channels i remember being subscribed to in my young zoomer teenage days. I don't intend be turned into a lolcow for schizos and spergs on these threads. The Internet and youtube back then was a lot more fun and welcoming to post content. Many youths around my age at the time were inspired by cobermani456 to start a youtube channel and post lets play content. Shame to see what has become of cobi and his mental decline especially after etika's death. Don't want to end up like him.
>>109393441
what about kandinsky labs?
https://huggingface.co/kandinskylab
https://github.com/kandinskylab
>>
>>109393513
Interesting notion. My only point is right now the newest big generalist model doesn't replace the specialized anime model we already have. Only time will tell if we'll ever stop needing one anime model and another for everything else.
>>
>tfw nochekaiser881 is now fully embracing krea2 and making anime character loras for it.

krea2 is finally picking of steam :)
>>
>>109393417
>specialized models will always beat generalist models
i mean ai itself disproves this as its a generalist thing that beats many specialized tools.

but yes, currently specialized models are better than generalist models, mostly because the generalists are not good enough. but once we have proper models that start generalizing across domains in a better way, it will be better than any specialist.
for example if we were to finally have a native multimodal image gen model with great language and image understanding, that would aid it in making better images.
>>
tfw literal who
>>
>>109392008
ego's
massive, massive ego's
>>
>>109392008
yeah i don't want anyone else to use my lora of my pp
>>
File: 257050011499651.png (2.43 MB, 1152x1536)
2.43 MB PNG
>>
>>109392008
They've never felt the joy of sharing a catbox with metadata and then days later seeing anon reshare it like "an anon posted this really good one awhile ago try it out it worked for me"

A single roach grifter named Furk stoled and then paywalled anons training settings nearly two years ago and ever since anons been schizo about it
>>
>>109393614
Why make separate lora for every single character? You could fit all Bleach characters in one
>>
>>109393724
Because Krea 2 can't stack loras.
ZIT can
>>
>>109393732
this makes no sense in this context
>>
>>109393740
Yes, it does. You can't stack multiple loras. That's why he's putting them out one by one. He's hindered by Krea 2's faggotry.
>>
>>109393753
>you can't stack multiple loras that's why he is releasing them one by one so you have to stack them
your stupidity knows no bounds
>>
>>109393763
You're just jealous because Krea 2 can't do porn
>>
>>109393740
if you didnt realize ur talking to a troll namefag falseflagger
>>
File: 192270917165743.png (2.07 MB, 1536x1152)
2.07 MB PNG
>>109392008
It be like that sometimes.
>>
>>109393770
>falseflagger
now THAT is a falseflag
>>
>>109393785
im not posing as anyone else, you dont seem to know what falseflag even is, "ZitChad" falseflagging retard
>>
>>109393795
confirmed
>>
>>109393724
all in one multiple characters lora packs are always bad and details either bleed or are inadequate per character tag. almost if not all of them are lazily done, terribly captioned and rushed out unbaked slop.
>>
>>109393770
Lies. I'm just a ZITGod looking to try to get Krea 2 BANNED for being a shit SFW model
>>
>>109393821
Sounds perfect for Krea 2 then.
>>
>>109393821
>all in one multiple characters lora packs are always bad
>almost if not all of them are lazily done
hmmmmmmmmmmmmm
>>
>>109393845
Sounds like a Krea 2 skill issue to me.
>>
>>109393655
If it's cute please share it
There's a sever lack of cute pp loras
>>
Saars.
>>
>>109393852
I agree brother. Fucking shit.
>>
>>109393645
Ego's what? What does he have that's so massive?
>>
>>109393858
Krea 2 is saved
>>
File: 67849589145833.png (2.12 MB, 1536x1152)
2.12 MB PNG
>>
File: ComfyUI_temp_yhsbp_00001_.png (2.82 MB, 1152x1536)
2.82 MB PNG
>>
woke up from a 2 week coma, has SDXL been dethroned yet
>>
>>109393920
Yes. By ZIT. In November
>>
>>109393920
we went back to SD1.4



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.