[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109844964

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/neo_collage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
File: shaboing-boing.jpg (642 KB, 1152x1792)
642 KB JPG
>>
>>
File: storybuk.jpg (600 KB, 1152x1792)
600 KB JPG
>>109850879
>>
>Working on a game
>Want to make some character sprites
>Generate a base character
>Realize local edit is so far behind, can't even say "put this character in this skimpy outfit while maintaining the style" properly
>Meanwhile GPT Image 2.5 can take 4+ character refs, put them all in a scene, turn them into pixel art, and dress them in unique outfits in a single prompt
The death of local edit HAS to be intentional at this point. Not a single model since shitty slopped Klein has even half-decent edit capabilities and it's been about a year. Why is local image so far behind? Somehow video is farther ahead despite being way more expensive
>>
Blessed thread of frenship
>>
>>109850879
Post workflow
>>
File: moonlight-screamata.jpg (700 KB, 1472x2176)
700 KB JPG
>>109850900
I suppose it's gimped on purpose to ward off 'deepfakes' or other unwanted uses, probably due to fear of lawsuits.

Let's wait a couple more months/years, hopefully we get fed some good local edit scraps soon.
>>
I will not be baited by Fear Uncertainty and Doubt posting
>>
>>109850909
nothing fancy just default + loras
>>
>>109850900
shame gpt butchers images with texture frying and inability to not touch anything it's not editing
>>
File: 4chan.jpg (429 KB, 1920x1088)
429 KB JPG
>>
File: ComfyUI_temp_tyeqc_00019_.png (3.15 MB, 1088x1936)
3.15 MB PNG
>>
>>109850946
I remember qwen-tts being alright
idk if there's anything newer
>>
File: geralt of feetia.jpg (424 KB, 1664x960)
424 KB JPG
>>
>>109851091
Looks neat :)

>>109851123
his foot is stuck in the dash, and humans don't look like that.
>>
>>109850857
Nice collage OP
>>
>>109851220
those threads cant have quality discussion with sabotaged ops. remove off-topic links from op and we can have cozy breads again
>>
>>109851389
This
>>
Continue crying :]
>>
cozy breas
>>
>>109850900
Image editing is something where I'm really considering just doing that with a cloud model, because image editing locally is seemingly years behind.
>>
File: ComfyUI_temp_dezel_00001_.png (3.04 MB, 1088x1936)
3.04 MB PNG
>>109851147
Thanks
>>
>>109851399
>>109851389
Can you prove these claims because the first thread had links looking in the archive. Are you claiming catjack created this thread and that anon decided to follow his lead?
>>
>>109851516
Answer the question schizo
>>
>>109850900
Skill issue
>>
>>109851540
Meds
>>
>>109851540
I see anons responding positively to his post....
>>
Is there anything new to generate pics with a transparent background with modern models?
LayerDiffuse only supports SD1.5 and SDXL.
Rembg and such totally suck, especially for things like hair.
I tried generating with a newer model and using LayerDiffuse and SDXL to extract an object but the result is not very good.
I saw that a Flux1 version of LayerDiffuse is being made, but I never liked Flux1. Hopefully they'll support Flux 2 Klein 9B.
>>
File: ss-lass_00031_.png (1.47 MB, 1024x1536)
1.47 MB PNG
>>
>>109850900
because the competition among video models is fiercer today. minimax is now very popular thanks to h3
>>
>>109850900
Local is just BFL. There's no other competent corp innovating with image models. H3 stole BFL's lunch so this time around either they're taking their time refining Flux 3 or we're not getting the model.
>>
>>109851729
Also the Chinese had so many different models, but yeah never released a single edit model on par with NBP. For some reason they all have the same idea of keeping their edit model behind API.
>>
>>109851729
I think it's because they realize that while the image model doesn't have as much, the edit model has commercial value.
>>
>>109851619
Use SegmentAnything to create masks instead. I don't know if it works with cumUI these days though
>>
File: Krea2_turbo_04290_.jpg (1.66 MB, 1776x2368)
1.66 MB JPG
>>
In H3, how do you use an environmental reference properly?

Even when I mark mine a weak reference, H3 has a habit of using the exact same camera angle and composition of the environment image ref, usually at the start of the video.

How do I make it so H3 only gets an idea that this is the location, but to avoid using it as the first frame anchor?
>>
>>109851729
>>109851736
a new version of qwen featuring the editing function is set to be released. just keep an eye out for when Kijai adds it
>>
it's a roach calm down, it's a miracle he grew fingers
>>
File: ComfyUI_06926_.png (1.12 MB, 1024x1024)
1.12 MB PNG
>>
>>109851767
visual mark the location. Like, draw a red box on your reference image and tell h3 this red box is the starting location of the first shot.

Also, it's not need but it's better to identity your location, <subject 2 > is the eviroment in < picture 2> etc...
Then H3 will keep surrounding area consistently
>>
>>109851834
>visual mark the location. Like, draw a red box on your reference image and tell h3 this red box is the starting location of the first shot.
Does that really work? Like telling the camera to start in a red circle? Example? Don't you have to specify rotation too? Seems like a bit of trouble.

>Also, it's not need but it's better to identity your location, <subject 2 > is the eviroment in < picture 2> etc...
Already do this. It doesn't prevent H3 from having a strong tendency to use environment refs as a composition anchor.
I also tried genning a video using 4 different environment refs: front view, right view, rear view and left view. The output video kept doing a 360 camera rotation to try and get all those individual references in, even though I asked for no such 360 panning shot.
>>
>>109850857
>mfw Resource news

09/19/2026

>Perceptual Refinement of an End-to-End Video Streaming Pipeline via Generative AI Layers
https://github.com/emanuele-artioli/presley

>PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance
https://github.com/StudioPiLabs/pace-core

>LingBot-World 2.0 realtime
https://github.com/kaarelkaarelson/lingbot-world-v2-realtime

09/17/2026

>Minimax H3 Latent Upscaler (3D) — BF16 Conservative v5
https://huggingface.co/Asirus/Minimax-H3-Latent-Upscaler-BF16-MAXQUALITY

>Kijai MiniMaxH3 INT8 VAE
https://huggingface.co/Comfy-Org/MiniMax-H3/commits/main/vae/minimax_h3_video_vae_int8_convrot.safetensors

>Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
https://zing.loopit.me

>comfyui-SelfLift: Progressive-resolution sampling for ComfyUI
https://github.com/facok/comfyui-SelfLift

>LTX-2.5-uncensored-v1.1-FP8
https://huggingface.co/frtertaer/LTX-2.5-uncensored-v1.1-FP8

>Introducing CUDA Rust: Two Tracks for Writing GPU Kernels
https://developer.nvidia.com/blog/introducing-cuda-rust-two-tracks-for-writing-gpu-kernels

09/16/2026

>Fizgig 6.0 - RefMod Training, Editing and Exploring
https://github.com/shootthesound/Fizgig/releases/tag/v6.0.1

>SlotDiT: Object-Centric Representations for Diffusion Transformers
https://slot-dit.github.io

>FastVideo-FastH3-8-Step-V2
https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2

>FastVideo FastH3 ComfyUI Repack
https://huggingface.co/FastVideo/FastVideo-FastH3-Comfy

>FastH3 V2 GGUFs
https://huggingface.co/realrebelai/FastH3-V2_GGUFs

>ComfyUI v0.36.0
https://github.com/Comfy-Org/ComfyUI/releases/tag/v0.36.0

09/15/2026

>LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
https://github.com/LynnReal-AI/LynnReal-Omni

>Meridian: Geometry-guided video model for authoring new observations of existing events
https://huggingface.co/Viggle/Meridian
>>
File: Krea2_turbo_03959_.jpg (1.46 MB, 2368x1776)
1.46 MB JPG
>>109851866
>>
>mfw Research news

09/19/2026

>Paint-Anything: Unified Any-Color Control for Image Generation and Editing
https://arxiv.org/abs/2609.20816

>Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network
https://arxiv.org/abs/2609.20633

>Printing the Underdetermined: Materializing Multi-solutionness in Figurative Paintings
https://arxiv.org/abs/2609.19782

>Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation
https://arxiv.org/abs/2609.19729

>FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants
https://arxiv.org/abs/2609.20769

>Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation
https://arxiv.org/abs/2609.19702

>Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models
https://arxiv.org/abs/2609.20139

>Astronex-World 1.0: Real-Time Interactive World Model Foundation
https://arxiv.org/abs/2609.20034

>VideoResearcher: Self-Improving Tool Design for Long-Video Understanding
https://arxiv.org/abs/2609.19664

>MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration
https://arxiv.org/abs/2609.19683

>Can 4D Foundation Models Remember?
https://arxiv.org/abs/2609.20819

>A Smaller Transformer in Your Transformer
https://arxiv.org/abs/2609.20100

>Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
https://choi-yeeun.github.io/MERIT
>>
Hit dog holler, ect
>>
>>109851866
>>109851869
Fuck off unemployed loser
>>
>>109851540
can you try to respond again without sounding so upset?
>>
File: Krea2_turbo_02673_.jpg (1.01 MB, 1776x2368)
1.01 MB JPG
>>109851965
No clue, It's more telling when they identify with abstract things that don't share their likeness.
>>
>>109851998
Yeah that guy is absolutely melting down over your gens. Pretty funny desu.
>>
>>109852022
well/ would you save that dogshit?
>>
>>109852022
What I don't get is his reluctance to make something of his own that could be funny.
>>
>>109851765
Kekkd
>>
>>109851998
>>
>MiniMax-H3 8-step PDD Acc LoRAs
good/bad?
>>
>>109852243
trash
>>
>>109852236
cool
>>
File: zI_00008_.mp4 (3.91 MB, 1312x768)
3.91 MB
3.91 MB MP4
>>109851862
test it yourself nigga
retention_analysis
<Subject 1> appears only in [Shot 2] after the zoom is complete: fully_preserved - preserve the recognizable facial identity, stylized design, hairstyle, and overall appearance from <Picture 1>.
<Subject 2> appears in [Shot 1] and [Shot 2]: fully_preserved - preserve the recognizable environment, design, and original 3D video-game appearance from <Picture 2> throughout.
Only use the white box as a reference location; do not show it in the video.
detailed_description
The target video uses a high-end 3D video-game style, but <Subject 1> must retain her original design, stylized facial proportions, and recognizable appearance from <Picture 1>. Do not turn her into a photorealistic live-action character.

[Shot 1] The shot begins exactly as <Picture 2> but without the white box. Do not show the white box. <Subject 1> does not appear in this shot. A bird’s-eye view of the city.

[Shot 2] At 00:01.00, from the original bird’s-eye view, the camera performs a quick dolly zoom toward the white-box location. Only use the white box as a reference; do not show the white box. The zoom completes and the camera locks completely at the target location.

at 00:02.00 Only after the zoom has fully finished and the camera is completely static, <Subject 1> walk out, She then raises her hand and shows a V-sign to the camera. The final framing clearly shows a wide shot of the surrounding environment and a clear full-body view of <Subject 1>
>>
File: debo_hd_k2_00043_.png (2.54 MB, 1664x1069)
2.54 MB PNG
>>
What's the deal with the deleted posts?
>>
>>109852480
unironic schizo meltdown
>>
I'm taking a huge gamble and experimenting doing a body transformation male 2 female lora for minimax h3 using seedance 2.0 videos. I have no idea how to properly train a minimax h3 lora and guides for it are all over the place. Burning a lot of money for this, please experiment with the lora.
https://gofile.io/d/GaHdaF7F
>>
>>109852390
not same anon. i followed your advice, but the first frame is still there at the beginning... the scene works well; i just have that annoying first frame..
>>
>>109852500
do you use first frame last frame workflow?
it never happens with ref2va workflow unless you specifically ask for it.
>>
File: model space.jpg (115 KB, 944x578)
115 KB JPG
i havnt genned in a while, kind of gave up porn altogether. but i dont want to delete anything just in case
>>
>>109850857
Did SLA attention nodes on Minimax increase VRAM use ? I turned it off and now my gen is working although i got more gen times compared with SLA on
>>
File: 1783820998637507.jpg (609 KB, 1642x2400)
609 KB JPG
>>109850900
Try Micro bikinis on GPT.
Or use Pic related.
Good luck
>>
>>109852487
I tried this and it created mustard gas.
>>
Training artist style lora for Anima
225 pics - 3000 steps - 1.0 learning rate - Prodigy - 1024x1280

Is that ok.
>>
>>109852678
you got oom?
that node has memory leaks
>>
>>109852889
No wonder. I use the Plaguekind version of SLA attention nodes.
>>
File: ComfyUI_Krea_2_00483_.jpg (3.04 MB, 2048x2048)
3.04 MB JPG
>>109847705
For alternating male and female vocals it's just
[Verse 1 - Male]
...
[Verse 2 - Female]
...
[Chorus - Duet]

and so on, it also suffices to use just describe duet in the caption instead of doing that.

https://files.catbox.moe/b3lbnz.flac
https://files.catbox.moe/bpp7zh.flac
>>
File: debo_hd_k2_00047_.png (1.88 MB, 1664x1069)
1.88 MB PNG
>>
>>109852899
i think that node is supposed to be used in conjunction with a sla turbo lora, if you're not using one of those then don't bother.
>>
>>109852952
Well shit. I cant use it alongside 10eros model. Even with non turbo version i got body horrors
>>
Wow this thread is literally dead now.

What are some other non-pozzed local diffusion communities? Or where did most people migrate?
>>
>>109852961
>this thread is literally dead now
local is literally dead now
>where did most people migrate?
nowhere because local is dead
>>
File: ComfyUI_Krea_2_00484_.jpg (3.67 MB, 2048x2048)
3.67 MB JPG
>>109852937
Nevertheless, I recommend reading
https://github.com/ace-step/ACE-Step-1.5/blob/main/docs/en/Tutorial.md

It's mostly applicable to YuE 2 as well.
>>
>>109852978
Why local is dead ? Minimax is just out a month ago.
>>
>>109852959
in my limited testing and i found that 10erosHybridV5 (non-turbo version) and the lightx2v ref2v 4step turbo v0.1 or lightx2v ref2v 8step turbo v1.0 768p worked pretty well for ref2video (euler). I didn't test i2v or t2v, so i can't speak to that.
>>
File: ComfyUI_Krea_2_00485_.jpg (3.27 MB, 2048x2048)
3.27 MB JPG
>>109852937
Hans Zimmer type prompt because why not?

Two takes with same prompt and lyrics
https://files.catbox.moe/rdbtsn.flac
https://files.catbox.moe/xrq8b8.flac
>>
File: ComfyUI_Krea_2_00488_.jpg (3.45 MB, 2048x2048)
3.45 MB JPG
>>109853132
Ah yes, Marricone style Spaghetti Western
https://files.catbox.moe/44x9w5.flac
>>
>>109850600
Krea2 still has no balls deep and blender viewport loras.
>>
>>109851866
>>109851869
thanks!
>>
what does new kjai int8 vae do?
i don't see anything
>>
>>109853194
Have yet to see a genre this model can't do, it's so refreshing
>>
File: Untitled.jpg (27 KB, 822x70)
27 KB JPG
>>109853259
supposedly also lowers vram usage
>>
>>109853264
Can Krea2 do porn?
>>
another day another vramlet crying about krea
>>
>>
>>109852961
First time being here between major model releases? Ah, I remember my first time too.
>>
Hibernation Mode: Ultimate
>>
Is there an image model that lets me use reference images, like sheets, styles, and generate completely new images?
>>
>>109853444
Techincally Minimax can do that.
But just wait for their proper image models
>>
>H3
This model is only good if you have an exact frame as ref, everything else is fine if you are just after some random slop.
>>
>>109853533
That's why a good image editing model is sorely needed.
>>
>>109853444
no. proper edit is always just 3 weeks away
>qwen 3
>flux 3
>krea 3
>h3 image
>>
>>109853344
hello
>>
>>109850857
this yue is proper slop
sloppiest of all ai slop

yet comfy made cover mode ability for it and not for ace step which has quality and antisloppering built in it.

how does one control bpm and key in yue slop?
>>
man, these 3d latent upscaler models are really shit
it's barely an upscaler. More like a low sigma denoise and retouch lmao.
I'd rather gen at native resolution.
>>
comfy just bends to the highest bidder
>>
When inpainting krea 2 images how should I prompt?
should I tell it "remove/replace item" or what? and what if I want to add another character interacting with the one already in the picture?
>>
>>109853643
ive never heard of a 3d latent upscaler
i still need to rape comfy with the aliased latent upscale desu
>>
>>109853444
if you had asked this in january anon wouldve said klein
>>
File: llada.jpg (3.29 MB, 6536x2802)
3.29 MB JPG
https://github.com/inclusionAI/LLaDA-Image
Has anon tried this?
>>
>>109853639
BPM and key are controlled directly in the caption. ACEStep has a habit of skipping lyrics and way worse genre comprehension so it's quite behind compared to YuE. Unfortunately, ACEStep XL might be the last we see of ACE Studio. They did not release their latest model (StepAudio 3), but their HF demo (https://huggingface.co/spaces/stepfun-ai/StepAudio-3-Music) and paper are out.

Technical report https://arxiv.org/html/2609.12945v1

It's a step up in audio quality, it's now about on par with Stable Audio 3 Medium, but this can do vocals. Its genre comprehension is still not as good as YuE 2 or Suno, but its out of the box sound quality is superior to both. It's a shame that once they got good, they decided to go closed.

Here's a few gens out of it. I would say YuE 2 is still much better in output quality and comprehension, bar its instrumental comprehension compared to this model. First output is 192 kbps so slightly was, rest are 320 kbps.


https://files.catbox.moe/u4c7ck.mp3
https://files.catbox.moe/gq12n5.mp3
https://files.catbox.moe/ya5mw8.mp3
https://files.catbox.moe/zqvxjw.mp3

This model is also censored. I would say YuE still has a good chance of coming out on top with a proper LoRA once they release the encoder.

Someone shared this melodic dubstep YuE 2 gen on Discord
https://files.catbox.moe/oyw7h5.mp3

This is what is possible with the community made encoder, which is slightly worse than official. So I'll look into training LoRAs with it for now.
>>
yue is like acestep turbo but sloppier
yet it has somewhat intradasting prompt interpretation
>>>/wsg/6237321

tho it is slop it can mimic live music
genre understanding is meme tier
but the good bit is that it sometimes but rarely adheres to [markers] as you specify them e.g. this song had
[Intro - song motif melody]

[Verse 1 - sombre and soulful, building tempo]
[Clean echoing electric guitar, slow melodic bass line, atmospheric and brooding]
All the leafs are brown
(All the leafs are brown)
Canada is ghay
(Canada is ghaaaaay)
I've been for a walk
(I've been for a waaaaalk)
on a pooo-poo day
(pooo-infested Canadian day)

[Verse 2 - sombre, building tempo]
[Clean echoing electric guitar, slow melodic bass line, atmospheric and brooding]
I'd be safe and warm
(I'd be safe and warm)
if I flew far away
(they would poo-poo far away)
Maple syrup dreamin'
(Maple syrup dreamin')
on such a winter's day
(on such a winter's day)

and it missed all of them.
some elements of song 'quality' come due to this example is obviously a 'cover' of a song
yue is slop.
>>
>>109853863
>melodic dubstep YuE 2 gen
YuE 2 LoRA gen*
>>
>>109853863
>slightly was
worse*
>>
>>109853863
>https://files.catbox.moe/ya5mw8.mp3
>https://files.catbox.moe/zqvxjw.mp3
those two are good altho second one is a bit silly others are slop-meh tunes


ty for the bpm info i figured it out by now
staying with ace step and have faith in them, no matter how long it takes them to release next model
>>
>>109853881
>how long it takes them to release next mode
It's highly unlikely unfortunately. They pulled a Wan on us and took the cloud pill. YuE 2 output quality is much more coherent than ACEStep XL without LoRAs, and with LoRAs it's only behind with instruments (though YuE 2 LoRAs are ahead of ACEStep). Not sure why you're still on ACEStep. Though ACEStep is still useful for LoRAs because the official YuE 2 encoder is not out, but I trust the YuE 2 to eventually release it alongside their technical report (as the dev has promised).
>>
>>109850857
>Music Models
Change to Audio Models
Sound cleaner
>>
File: 1764112691215358.png (1.37 MB, 832x1216)
1.37 MB PNG
>>109853741
>vintage photograph of a very fat bald and ugly man sitting at his dirty desk in his dirty room. his hands are on the keyboard and the computer monitor reads "local diffusion general". he is looking at the monitor while the viewer is positioned behind the man, flash photography
>>
>>109853682
No, i think inpainting is the same as it has been in stable diffusion. There is no actual edit functionality.
>>
>>109853639
Also a tip for getting better non-bland gens out of YuE 2, set the CFG to 1.2 or higher. Pretty much my default.
>>
File: 1773673455371157.png (1.45 MB, 832x1216)
1.45 MB PNG
>>109853977
>Vintage flash photograph, shot from directly behind a very fat bald man seated at a cluttered desk in a dirty, dim room. The camera looks over his bare shoulders and the back of his head; his hands rest on a beige keyboard, and the CRT monitor in front of him clearly displays the white text "local diffusion general" on a dark screen. Grimy walls, stained carpet, scattered cans, papers and cables around the desk. Harsh on-camera flash blows out his skin and the nearest surfaces, hard shadow of his body cast onto the wall behind the monitor, deep falloff into the corners. 35mm film grain, slight color shift and mild lens vignetting, amateur snapshot framing, slightly crooked horizon.
>>
>>109853984
thx for the info will test
did try those lora links posted yesterday, instrumental one is good but those other comfy ones make zero difference
>>
File: 1763328826612912.png (1.26 MB, 896x1120)
1.26 MB PNG
Also does editing
>make her breasts larger
>>
File: 1765342214208956.png (1.96 MB, 1216x832)
1.96 MB PNG
>>
sulfur when
>>
>>109853993
Can you try a candid photo in a busy sidewalk? It almost looks good but the examples in the preview seem to be lacking details in the background
>>
File: 1769434934375087.jpg (639 KB, 2496x1216)
639 KB JPG
>>109854156
I think it was only trained at 1024px, on a different gen it butchered pupils. It also takes ages for the MoE encoder stage to finish on my card and the authors sampling defaults are trash. It might have some nice aesthetic pockets in its dataset but overall it's pretty meh at first glace.
>>
File: 1779322659301704.png (1.31 MB, 832x1216)
1.31 MB PNG
>>109853993
Same prompt but with my usual sampling params instead of the authors' default. They recommend euler and their custom llada_image scheduler. My own are much better looking but still lacks.
>>
>>109854217 (You)
I take that back, the model with my params looks equally as shitty but in a different way.
>>
>>109854202
Thanks, yeah it seems severely underbaked
>>
File: 103258CUI_00001_.png (543 KB, 1216x832)
543 KB PNG
>>
>>109852961
The schizo got bopped and is having a harder time now that he's a confirmed praig that actually wanted to bottom for comfy. Comfy wasn't bisexual and it took him over the edge.
>>
>>109854612
ran can you fuck off with your cringe faggot fanfiction already and let us discuss local diffusion in peace
>>
>>109854639
You sure about that bud?
We can all read OP
Usecase for posting this after a self dox?
>>
local diffusion?
>>
>>109854639
Totally normal things for a man in his mid thirties looking to get a job and support for a project to post even on /d/ of all places.
Is he not aware that doing stuff like this has consequences?
I'm afraid scroll down after reading this post, his post history is fucking disgusting. But one thing is certain is that he loves dick.
>>
>>109854703
>>109854736
can you fuck off with your schizo bullshit nobody wants to see this retarded discord drama here
>>
>>109854736
This is what kids call "rent-free". My God.
>>
The anon seething on behalf of ani is reacting in the same manner as the schizo that got banned last night. Who we can confirm was samefagging.
Also who is ran?
>>
File: AniStudio-95632.mp4 (3.85 MB, 768x1376)
3.85 MB
3.85 MB MP4
>>
The newer Int8 Video Vae is just a placebo
>>
can i run krea with 12gb vram and 16gb ram?
>>
File: 00006-1061641.jpg (679 KB, 1728x2880)
679 KB JPG
>>109852961
with the constant negativity, shitposting and schizo posting. what did you honestly expect would happen to this place? The only people getting (you)s are shitposters and "debo" larpers and attackers. A bunch of made up characters and personalities draining up the energy and attention in these threads. Constant shitty collages and thread bakes made genuine serious people disinterested in posting here.
>>
well
i must say that this yue might be sloppy but it aint too shabby

my ode to canada
nugen version
>>>/wsg/6237410
>>
>>109854905
Current collage looks fine to me.
>>
>>109854048
those breasts look really bad
>>
I'm at step 3,000/6000 added another epoch to the gofile.
https://gofile.io/d/GaHdaF7F
>>
>>109854922
Just always feels a bit lazy when it doesn't have any videos.
>>
>>109854905
its still mostly anima shills fucking shit up here. cant do anything about it, those bots will suppress any genuine discussion and criticism of cumfart and their garbage models
>>
>>109855096
Nobody is discussing anima
>>
>>109855096
apache 2 tranima status?
>>
>>109855113
yes but they are more subtle than they used to be
>>
>>109855121
All I see is ani and his noobcord friends crying over not getting funding. Nobody cares krea is the future. Looking at the above post I see why he will never get support or funding
>>
>>109855139
rent-free
>>
>>109854967
https://civitaiarchive.com/models/2942542?modelVersionId=3331615
>>
File: 00014-88045382.jpg (648 KB, 2880x1920)
648 KB JPG
>>109855096
i think its just cynical and demoralized vramlets and poorfags in general that are shitting up these threads.
it has to a painful to not even be able to run krea2 at nvfp4. Day by day I'm seeing more interesting krea2 loras get posted.
https://civitai.red/models/2804299/krea2loralab?modelVersionId=3339168
someone even made minimax h3 lora of re:zero
https://civitai.red/models/2949009/rezero-anime-style-character-lora-h3?modelVersionId=3339742

>>109854890
possibly but you serious need more ram, like double amount of that for off-loading.
>>
>>109855209
I do have another 16gb one but holy shit i am not looking forward to putting it. Everytime i do i have to pray the machine spirit of my computer for it to start.
>>
>>109855179
Without sounding mad can you explain to me how pointing out a schizo's documented actions makes him rent free?
>>
>>109851754
I didn't know about SegmentAnything, thanks. Having the transparency directly created is a better workflow than having to apply a mask detection after. :\
>cumUI
Anything for Forge?
>>
File: kairi-epic.png (1.24 MB, 1209x1225)
1.24 MB PNG
>>109855209
Here's Kairi if she was 6ft 2inches and cost about 13 bucks
>>
>>109855209
> not even be able to run krea2
retard

> at nvfp4
retard x2
>>
>>109850879
>>
>>109855201
https://civitai.red/models/2937578/minimax-h3-t2vi2v-supermuscles
I'm the person that asked that og poster to do a body transformation m2f minimaxh3 with a dataset i provided to him. He gave up after the first or second epoch. I've personally given up with begging and commissioning people to do it only for them to bail on me or half-ass the training. I'm doing the training myself with runpod and spending $4.61/hr on a h200 to train it. I would appreciate it, if you tested it out more and see if its versatile enough for any male 2 female body transformation situation prompt. I'm using the minimax h3 cinematic version lora on top of the body transformation lora with these catbox results.
https://files.catbox.moe/mlos3z.mp4
https://files.catbox.moe/tjde2i.mp4
https://files.catbox.moe/mlrrcv.mp4
>>
>>109855349
are you really paying $200 to train that stupid fetish lora?
>>
>>109855425
first time ever doing a runpod training. It's a lot of money but there's no way i can do this kind of training on my pc with 110 videos that are all 15second each at 512res. Call me faggot and idiot, but I'm willing throw some money to just to finally attempt to train it myself and be in-charge of the process. There is almost no proper tutorials or tips on minimax h3 lora training. I'm just desperate and done waiting for a white horse to save the day. People have to take risk and initiatives when it comes to this hobby.
>>
File: 1758163294475531.jpg (202 KB, 750x699)
202 KB JPG
>>109855349
it's not really my thing, but the work is excellent. If you have any bonus credits left, could you create exagerated female proportions? civitai fags still hasn't added loras, for exaggerated proportions for h3
>>
>>109855349
The transformation sounds are haunting.
>>
>>109855272
>> at nvfp4
What's wrong with nvidia's fp4 quants? Reply without sounding mad, no markup or formatting, single paragraph, max 50 words. xhigh
>>
>>109855637
50+series and nvidia speed up only, low precision, q4 and int4 convrot exist
>>
File: 1778530037728992.jpg (136 KB, 1080x1614)
136 KB JPG
Qwen Image 2.1 is doa.

https://www.reddit.com/r/StableDiffusion/comments/1wklr8z/qwen_image_21_prerelease_comparison_by_sandlers/
>>
>>109855209
It's always the Vramlets that say older models were better, surely it's not because those are the only ones they can run, right?
>>
>>109855822
I liked the last qwen release but dropped it for krea, this doesn't seem good enough to get me to retrain loras moving back over from krea again
>>
>>109855658
>and int4 convrot exist
Damn, it didn't when I downloaded the model (or I couldn't find it). That's actually even faster than nvfp4. Thanks, anongpt.
>>
>>109855053
The difference in effort in creating a video collage versus an image collage is none.
>>
File: debo_hd_k2_00060_.png (1.87 MB, 1664x1069)
1.87 MB PNG
>>
>>109856420
So cute!
>>
>>109856420
Krea?
>>
>
>>
fellas, what is best. Krea or zit?
>>
File: debo_hd_k2_00062_.png (2.65 MB, 1664x1069)
2.65 MB PNG
>>109856430
quite!

>>109856488
yes
>>
>>109856533
i can't remember the last time i saw anyone comment or use zit. i think everyone dropped it after krea2 released
>>
>>109856640
woah, is it that good? What checkpoint you guys use?
>>
File: 1780960439328180.png (1.14 MB, 1048x1058)
1.14 MB PNG
Extra credit!
https://files.catbox.moe/xm10yw.png
https://files.catbox.moe/zyh6be.png
https://files.catbox.moe/r01nl8.png
>>
>>109856641
i think for anime you use anima (or an anima finetune) and for everything else you use krea2 (in terms of image)
>>
>>109856640
copejeets with 8gm vram cards sometimes use it, but if you can use krea there is no reason not to
>>
>>109856600
Why don't you go back to your stable diffusion only general? Bored of the same 2 "anons" spamming all day there?
>>
>>109856646
Nice
>>
>>109856743
kys catjak
>>
>>109856648
Krea is also pretty good at anime, if you're not looking to just generate a specific 1girl.
>>
>>109856848
>>109855301
>>
>>109856646
Now make her a teenage celeb plox
>>
>>109856848
i want specific artstyles, not for me desu
>>
>>109856860
glow harder
>>
whats the ideal quant for krea2 on a 12gb vram?
>>
>>109856884
int8
>>
File: file.png (6 KB, 405x98)
6 KB PNG
>>109856948
but hardening is bigger than 12gb...?
>>
>>109856952
it's ok
>>
>>109856648
>for anime you use anima
Nope! Even anons on /adt/ prefer krea2 over tranima. It's a failed model with zero finetunes so it's best to forget about it
>>
>>109856959
Really? How? I thought image models needed to fit it whole into the gpu...?
>>
>>109856963
nah, anima is good
>>
>>109856963
>>109856520
>>
>>109856850
Is that with a lora? It knows certain 1girls, but not many.

>>109856861
It knows lots of artstyles. If you just mean the specific style of certain pixiv artists (like ATDAN etc.), it can't help you there. That's where Anima is better.
>>
>>109856967
bruh
>>
>>109856979
?
>>
>Even anons on /adt/ prefer krea2
adt is like a joke tho dont you know anon?
>>
>>109856974
Thanks for alerting me to a new CatTower being out, even if that wasn't your intention.
>>
>>109856997
90% of the posts there are ai generated style not actual anime heh
>>
>>109856997
Krea2 is slopped to hell and back for anime, though it's nice as img2img 2-pass for Anima (first gen on Anima, then second gen on Krea).
>>
I get this thread is mostly for visual diffusion models, but has anybody done any work with text diffusion models? I've got SFT and DPO up and running for my text diffusion model, but I need some way to excise the corpo-slop tendencies from it. J-wash was great for the autoregressive model but abliterating the text diffusion model causes problems with the output and regresses key parts of the DPO phase. I've spent the past week and most of my budget with two different coding assistants trying to figure out how to do any form of impactful abliteration or even feature steering and I remain completely dissatisfied with the results. There are a handful of papers about ablating concepts in image models but absolutely nothing about text models. Anybody have any advice in this thread?
>>
H3 image, krea 3, flux 3, qwen 2.1
and nothing ever happens
>>
>anime
You're either using NovelAI or Anima.
>>
>>109857068
ldg is where horny fat guys hang out and generate nude 1girl. the smart people hang out on lmg; they'll answer your questions
>>
anima is mega dead, it just got parameter-mogged by krea2.
>>
>>109856969
Good is where Ill is at and it isn't really better than that. Krea2 is the best

>>109857024
That is a skill issue and shows you don't know any anime shows or styles at all
>>
>>109856993
>>109856979
come on dude, just tell me if youre bullshitting or if i need to get a lower quant lol
>>
>>109856997
>>109857100
wow thanks for the info you have totally convinced me to give you a million dollars to make an anima copy but based on krea
wait no i'm not a fucking retard
>>
>>109857077
We'll see. Minimax H3 pretty much came out of nowhere. Felt that way to me, at least. There seems to be a great reluctance to release models capable of editing images, however.
>>
>>109857131
we only need 1 good model. before we had anima, we were stuck with garbage image gen too. if we get 1 good image edit model, it's over for the cloud jew as far as image gen is concerned
>>
>>109857110
cumfart can split weights
think it's like two models and you can load one, do job, unload, and do the same with the second
>>
>>109857131
They're probably just much more difficult to train than standard image models, since they have to actually understand and follow instructions in natural language and there's not much freely available data for image editing tasks in the first place. Not even MiniMax H3 is that good with following instructions outside of standard use cases.
>>
File: GabagoolStudio-95632.mp4 (3.3 MB, 768x1344)
3.3 MB
3.3 MB MP4
>>109854824
>>
File: ComfyUI_07974.jpg (1.64 MB, 1200x1800)
1.64 MB JPG
>>109854048
Add "oiled breasts" to the negs, it usually helps.

>>109855822
Is it as compute-heavy as the last model? All I remember about my (admittedly brief) time with a Qwen image model was the slow image generation/editing. Also, that guy's test images are pretty bad all around. So bad that they tell me nothing about it's actual capabilities.
>>
>>109857308
I found Minimax H3 to be fairly competent at putting reference characters into scenes. I don't need much more than that, and an image model that can just spend all its compute on a single image shouldn't suffer from the slight quality issues H3 exhibits when trying that, I imagine.

>>109857421
You'd feel a lot better, if you took them.
>>
>>109856884
Just go Q4/FP4/INT4/anything4
There is no real perceptible quality difference between 4, 8, and 16.
>>
Last 2 months I've been living under a rock.
Any major news?
Because at the first glance things seem rather quiet.
>>
>>109857734
H3 and Anima
>>
>write 1 minute long prompt
>set duration to 5seconds
kino enhanced
>>
>>109857308
>They're probably just much more difficult to train than standard image models, since they have to actually understand and follow instructions in natural language and there's not much freely available data for image editing tasks in the first place.

I don't think ClosedAI have changed their training methods. It's all AI, so they have a model capable of making edits and describing those edits, or similar.
>>
>>109857734
local is stronger than ever
>>
File: SydneyComparison.jpg (3.72 MB, 8800x1344)
3.72 MB JPG
Using a Lora of a someone even if it's just a regular non-edit Lora helps a lot with fine details when upscaling other photos of them that weren't in the original Lora dataset, with Klein 9B (Distilled)
>>
comfortable breas
>>
Maximin H6
>>
File: failure.mp4 (535 KB, 1376x768)
535 KB
535 KB MP4
I need help bros. I'm a FAANG engineer. I bought a 5090 so I can develop my way out of this grind. But I realized I have no creativity and have no idea how to retire quick. What's the best get rich quick scheme with AI?
>>
Did anyone ever figure out what the "Z" in Z Image stands for? Or what "krea" even means?
>>
>>109858416
Just run it through DLSS5 if you want details.
>>
>>109858466
>hag filter
I'm not Indian
>>
>>109858460
thats why i got a 5070 ti. I'm creatively bankrupt so I'm waiting for a true VRAM use case rather than buying VRAM and

>now what

although I wish I had it now for Minimax H3
>>
>>109858486
Not really my problem, is it?
>>
>>109858460
the money isn't in AI content, its in AI software. if you were actually a FAANG eng, you'd already know this
>>
>>109858460
>What's the best get rich quick scheme with AI?
either sell gay pony porn
or finetune models so they can make gay pony porn
>>
>>109858460
as a FAGMAN you should be able to figure it out yourself
>>
>>109858460
make a krea 2 lora, merge it and sell it as a full anime finetune
>>
>>109858611
being an engineer doesn't mean i have the business acumen. why would i be sourcing opinions if I knew how to unblock myself
>>
>>109858636
not true. i and 90% of my teammates are dumb. maybe 10% of the company actually has a brain
>>
>>109858643
who tf is going to buy a lora
>>
>>109858460
we're in desperate need for a rust based UI that can run models like minimaxh3 directly. if you do that, with some kind of timeline view, you'll have made the king of UIs for 2026
>>
>>109857539
HE'S BACK!
>>
>>109858460
feels good to be a creativitychad. AI exposed all the frauds like you and I'm glad you're getting filtered.
Now I just need to wait until AI fixes laziness and then you'll see what I'm capable of.
>>
>>109858460
your example does not prove that you have a 5090
>>
>>109858466
how on earth would that be better than using an edit model + Lora trained on actual photos of the exact person
>>
>>109858643
kek
>>
>>109858460
gay / furry commissions
not joking
>>
>>109858465
Neither mean anything AFAIK
It's still funny that Z Image ISN'T a model made by Z.AI though lol
>>
>>109858818
You wanted to upscale and add detail. DLSS5 does that, but faster.
>>
uploaded the male 2 female body transformation lora to civitai, gave up at 4500/6000 steps. maybe next i'll try rank 32 and but reduce the amount of clips.
https://civitai.red/models/2950225/body-transformation-male-to-female-minimax-h3-fl2va?modelVersionId=3341226
>>
>>
>>
>>109859202
>>109859212
ill give it a pity download for your trans fetish but I am judging you
>>
>>109859243
At least he's open about his mental illness unlike
>>109854703
>>109854736
Who post embarrassing shit like this and tries to get it scrubbed
>>
>>109859322
When did ani try and scrub any of this? Do you have proof?
>>
Haven't used this stuff since like SD1.4 so I'm way out of the loop. Wondering what I should use if I want more creativity than consistency or even quality these days. Making a SW Battlefront knockoff and I want to give it some vague prompts and just hammer the GPU until I see something I like.
>>
>>109859500
She made boom boom and sperged too close to the sun and got 5 /adg/ threads deleted.
>>
>>109859721
krea2
>>
File: dt-lands.jpg (451 KB, 1664x960)
451 KB JPG
>>
File: drinkerlands.jpg (392 KB, 1664x960)
392 KB JPG
>>
File: poopjuice.jpg (557 KB, 2240x960)
557 KB JPG
>>
File: da-juice.jpg (522 KB, 2240x960)
522 KB JPG
>>
>>
any good LLMs (plus prompts for that) for helping proooompt with anima?
>>
is krea2 really better even for anime than anima?
>>
File: ComfyUI_01166_.png (1.19 MB, 1024x1024)
1.19 MB PNG
>>109859907
>>
>>109859937
no. it's just easier to do it with krea2 then just crank the attention on certain parts. anima is a failed model
>>
Can anons not treat the thread like it's an art share and actually discuss tech?
>>
>>109859953
What do you want to discuss?
>>
>>109859946
isnt krea2 censored to hell and back though?
>>
File: ComfyUI_01168_.png (1.29 MB, 1184x880)
1.29 MB PNG
>>109855209
>>
>>109859939
no
>>
haha anima is so good!
>gets 6 finger gens
>>
File: no-clace.jpg (917 KB, 1792x1792)
917 KB JPG
>>109859969
>>
>>109859963
>>109857641
>>
>>109859956
People that need an llm to prompt have no talent
>>
>>109860009
Depends on the model, I typically use it for formatting then modify it myself after
>>
>>109859963
anima is better for nsfw content. use krea if you want better combat and serious sfw content
>>
File: geralt-of-kartia.jpg (839 KB, 2048x1536)
839 KB JPG
>>
>>109860034
This is a blue board so that is irrelevant. Anima still mogs no matter what and it's easier to train loras for it whick makes the nsfw barrier trivially easy to bypass
>>
File: geralt-and-young-hilldawg.jpg (898 KB, 2048x1536)
898 KB JPG
>>
>>109860034
what checkpoint should i use? Or is everyone using base?
>>
>>109860076
>Anima still mogs
*Krea2*

>>109860081
Krea2 turbo
>>
>>109860177
krea requires too many loras for simple nsfw content. but the image quality is indeed superior
>>
>>109860205
Anima requires too many loras to be competent at all
>>
yes, hello?
is this the right way to the kinoplexatorium?
>>
File: 1764500100196.jpg (65 KB, 822x1024)
65 KB JPG
I'm about to work on an important lora for krea. Can any lorachads give advice on what scheduler and so on are best? I've been using automatic v2 so far, and everything else as the ostris default settings which seem to work OK most of the time.
>>
>>109860558
aitoolkit, all defaults, anything else I have tweaked with just makes it worse and it takes too long per lora at the size of krea for me to want to explore much more
>>
>>109858460
Make something better than Comfy and written in .cpp. Using Comfy is like using an Apple device. It's cool with Python, but cpp is better. There have been community attempts, but all of them kind of suck.
>>
>>109854890
I can run Krea at int8 with 16/16, but the vram fit is pretty exact, so 12GB might need to do some shuffling. Use --disable-smart-memory so it doesn't try holding more than one model in vram at a time. I think keeping dynamic vram enabled should help avoid pagefile writes too, or at least it's what I do.
>>
File: 0920104659123-i1xsVmSUjB3.jpg (476 KB, 2439x1371)
476 KB JPG
This model is so horny, I couldn't stop h3 from generating nipples/nipslip on scantily dressed anime girls. I tried prompting "no nipples", "safe for work", "covered breasts", they don't really work.
https://files.catbox.moe/e5a285.mp4
>>
>>109860773
not really, it's not 2024 anymore
we have great .cpp UIs
>>
>>109860939
not really. there is only one actual cpp application that isn't npm slop but it will make the thread schizo have a complete meltdown if I mention it
>>
>>109860890
Maybe it doesn't believe it's physically possible to stay up? What if you explained why no nipples, like her neckline is fashion taped to her breasts to hide nipples or something.
>>
>>109860890
My H3 defaults to very large boobs and childbearing hips.
You should not even mention the words in the prompt as there are no negatives in this model.
Also it doesn't seem to understand some words at all, probably because censored.
>>
>>109860890
i get this on 10euros all the time
>>
>>109858416
> eyelashes
no it sucks
>>
hellow anon
>>
File: that's a deep creek.mp4 (2.8 MB, 1280x720)
2.8 MB
2.8 MB MP4
>>
>>109860975
>there is only one actual cpp application that isn't npm slop
none of the UIs linked on the sdcpp repo are npm nor slop, retard
this is why every dev hates you and you will never be accepted in dev spaces, ani
>>
>>109861337
They are all ts or js retard. there is one gimp plugin but gimp isn't really a good UI to begin with
>>
>>109860890
>give horny reference
>get horny result
working as intended
>>
A local kinnosuir at the kinodrome told me to come to the kinolosseum to watch the kinovision
>>
>>109861337
Nvm, you are the subhuman that ruins the threads with your meltdowns. You are ban evading btw
>>
File: file.png (1.31 MB, 1280x960)
1.31 MB PNG
people dicksuck krea2 but i ask for subject 2 to tower over twice the height of subject 1 and thats what i get.

frankly disappointed lol
>>
>>109861442
Instead of being retarded you just describe the knight as being twice as tall as the girl then cranking the attention around those tokens. Same thing for getting big boobs or sexo. Weighting tokens is essentially the way to break the model to your whims. Also, post in the anime threads if that's all you make
>>
>>109861456
nah i just downloaded it and i am trying to understand how i use this shit before making any prompts that require effort.

And what you said is exactly what i've done. I guess krea just has shit spatial awareness.
>>
>>109861425
My kinosovl vision is activated I am seated
>>
>>109861463
If you are doing weak ass attention like :1.5 you are doing it wrong. It's a scale from 4.0-10.0. dit need MOAR
>>
>>109861442
post here, don't listen to that faggot
>>
BAKE?!?!!??!
>>
>>109861360
keep shitting on devs more successful than you, I'm sure it'll make you rich, even though you've been unsuccessful and unemployable for 2 years
imagine being such a talentless hack your only cope is smearing your competitors on an anonymous imageboard
i feel bad for hlky
>>
Anon cries out for my handle and cries out for my gens
>>
>>109861501
Make one without the schizo rentries. The actual schizo will just get himself banned again
>>
New:
>>109861523
>>109861523
>>109861523
>>
>>109861463
>i'm a retarded noob and i don't know what i'm doing
>i can't get the model to do what i want!
>the model must be at fault!
idiot
>>
>>109861533
>>109861476
Here you go anon. Ordering matters more too.
>>
>>109861518
Did that post sound convincing when you typed it out
>>
>>109858460
sell 5090s to idiots with no imagination



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.