[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Brief downtime for maintenance shortly. Maybe 10 minutes. Idk.

[Advertise on 4chan]


File: Krea2_turbo_03114_.jpg (1.54 MB, 1776x2368)
1.54 MB JPG
Discussion and Development of Local Image, Video, and Music Models

Previous:>>109686888

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

08/30/2026

>ComfyUI MiniMax H3 Prompt Writer
https://github.com/duckyshell/ComfyUI-MiniMaxH3-Prompt-Writer

>ComfyUI-HR-Endless-SamplerL Chunked video/audio sampling for ComfyUI with latent continuation
https://github.com/hradec/ComfyUI-HR-Endless-Sampler

>ComfyUI_MiniMax_H3_Extender: FL2VA support + major speed and memory optimizations
https://github.com/tritant/ComfyUI_MiniMax_H3_Extender/releases/tag/2.0.0

>MiniMax H3 Super Acceleration · ComfyUI 模型包
https://huggingface.co/t8star/Minimax-H3-Super-Acceleration-Comfy

08/29/2026

>Fizgig v5.0.0 - Full fine-tuning for Minimax and Krea 2 for 16gb+ VRAM
https://github.com/shootthesound/Fizgig/releases/tag/v5.0.0

>LoRA Dataset Studio: Complete, self-hosted LoRA workflow in one browser tab
https://github.com/perfectgf/lora-dataset-studio

>FL MiniMax H3: MiniMax H3 workflow nodes for ComfyUI
https://github.com/filliptm/ComfyUI-FL-MiniMaxH3

>Kijai / Minimax H3 Fastvideo vsa_datafree_1300step_4step_int8_convrot
https://huggingface.co/Kijai/MiniMax-H3-experimental/blob/main/minimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot.safetensors

>h3-storyboard: A Claude Code skill for turning a script into MiniMax H3 shot lists
https://github.com/phileiny/h3-storyboard-skill

>Suno AI's source data leaked
https://archive.org/download/suno-ai

08/28/2026

>FastVideo FastH3 V1: Open-Weight 4-Step Sparse Distilled Minimax H3 for 14x Speedup
https://haoailab.com/blogs/fasth3-preview

>ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI
https://rocm.blogs.amd.com/ecosystems-and-partners/rocm-x-blog/README.html

>H3 Prompt Composer V5.43.4
https://github.com/BMB12d3/minimax-h3-prompt-composer/releases/tag/V5.43.4

>PotionUI
https://github.com/PotionUI/PotionUI

>Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
https://wucy0519.github.io/MMLVE

>Comfy announces official resale of MiniMax commercial-use licenses
https://comfy.org/minimax/license
>>
>mfw Research news

08/30/2026

>GaussVid: Sparse-View Gaussian Splatting with 3D-Aware Video Diffusion Priors
https://arxiv.org/abs/2608.21849

>Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation
https://rex0191.github.io/DeMoDiff

>Adapting Dense Vision-Language Relationships for Multi-label Classification with Partial Label
https://arxiv.org/abs/2608.22313

>PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
https://arxiv.org/abs/2608.27345

>LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
https://arxiv.org/abs/2608.27395

>EchoWM: Open and Enterable Omnimodal World Models
https://arxiv.org/abs/2608.23189

>MoTE: Mixture of Task Experts for Multi-Task Video Understanding
https://arxiv.org/abs/2608.24763

>Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning?
https://arxiv.org/abs/2608.23074

>Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models
https://arxiv.org/abs/2608.27065

>Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning
https://arxiv.org/abs/2608.26993

>SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image
https://arxiv.org/abs/2608.23930

>Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
https://arxiv.org/abs/2608.26317

>LLaVAFlow: Preserving Latent Alignment Flow for Parameter-Efficient Multimodal Fine-Tuning
https://arxiv.org/abs/2608.26820

>Training-Free VLM Personalization via Calibrated Residual Decoding
https://arxiv.org/abs/2608.22263

>On Tensor-Based PDEs and their Corresponding Variational Formulations with Application to Color Image Denoising
https://arxiv.org/abs/2608.22302

>What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
https://arxiv.org/abs/2608.00013
>>
>>109689804
debo, nobody clicks your garbage. you're a cretin.
>>
>schizobake
check
>No collage
check
>news for the schizo to rage about
check
>samefagging the first few replies itt
>>109689815
>>109689822
>>109689827
triple check, to another great thread guys!
>>
Blessed thread of frenship
>>
>schizo
>schizo, schizo: "schizo, schizo schizo! schizo!"
>>
>>109689792
I look like her (I wish)
>>
>>109689889
Let's lay off of him, he's going to bait to derail the thread, His time is worthless but not yours
>>
File: debo_mcn_k2_00143_.png (3.88 MB, 1536x1280)
3.88 MB PNG
I haven't been around much this morning. why's he having a nuclear melty?
>>
Is this the fake bake?
>>
>>109689905
This clown really think he's beefing with a Hecatoncheires and can do 3+ post back to back
>>
File: 85012.png (95 KB, 1724x624)
95 KB PNG
MiniConstruct 0.2.0 is out (modular prompt builder for H3).

https://github.com/InsertSpice/MiniConstruct

Supports:
-Camera action toggles
-non_diegetic_music switch
-Visual style/Medium selection
-Subject identity fidelity
-Tone adjustment
>Reference roles
>Iterative selection editing/regeneration

If any of you want to try an H3 prompt builder, give it a shot.
>>
>>109689792
Is that a paddle or a hand mirror on the bed?
>>
>>109689905
cute
>>
File: Dunkirk.mp4 (3.94 MB, 1152x656)
3.94 MB
3.94 MB MP4
The Horror!

https://files.catbox.moe/pojg9y.mp4
>>
>>109690016
Paddle
>>
>>109690025
kek
>>
>>109689792
Purdy. Prompt? Wondering what kinda art style you specified, looks to my untrained eyes like 19th century Frenchies.

>>109689876
YWNBAW (and that's okay).
>>
>>109689905
Magical flora reclaiming a druid's abode, nice.
>>
>>109690043
I would share more but there's a small circle of salty posters that have followed me for years to discredit me. This includes impersonating me using any prompt they could get their hands on and you're hitting the nail on the head, krea 2 is good at grabbing time periods.
>>
can lilbro stop samefagging his own gens?
nb4
>>
need rentry for ran faggot
>>
File: Matrix.mp4 (3.28 MB, 1152x656)
3.28 MB
3.28 MB MP4
H3 kino accent alert

https://files.catbox.moe/2tc3nr.mp4
>>
>>109689970
still don't know what model I should use.
>>
>>109689792
thanks for baking this thread, anon
>>109689863
thanks for blessing this thread, anon
>>
>>109689804
>>109689813
thanks for baiting this thread, anon
>>109690080
thanks for thanking this thread, anon
>>
>>109690061
reminder to never share prompts as ai jeet grifters will try to monetize it right away
>>
File: 756586130238784.mp4 (3.79 MB, 928x704)
3.79 MB
3.79 MB MP4
>>
Open source your workflows. Secret prompt? C'
mon, man.
>>
>>109690074
lmao holy shit
>>
>try kroma v0.2
>works nicely, barely any body horror
>try kroma v0.3
>start getting missing limbs and bad anatomy on the same prompts/settings right away
LODESTONEEEEEEEES!!!!!!!!!!!!!!!!
>>
>>109690103
He can't keep getting away with it!
>>
>>109689533
it would be good if it worked
you write all that shit like picture weak reference transfer only subjects positions but it copies background shape and everything else
>>
>>109690103
is there an audio version?
>>
File: Return_00194_.jpg (1.41 MB, 1776x2368)
1.41 MB JPG
>>109690086
I can give the prompt but they will be disappointed because I use my own custom nodes that combine neg pip for krea and prompt to and from.
I don't trust most node creators and create my own
>>
>>109689804
>>MiniMax H3 Super Acceleration · ComfyUI 模型包
>https://huggingface.co/t8star/Minimax-H3-Super-Acceleration-Comfy
anyone try this yet?
>>
>>109690147
No
Only dumb newfags click random links the OG schizo post. Please read the renty to see why
>>
>>109690153
Fuck off. No one cares about your off-topic schizo drama. I just want my H3 gens to go faster.
>>
>>109690173
Then you're free to try it out "anon"
I can't wait to see what you make
>>
>>109690043
> YWNBAW
>that’s ok
It’s not ok anon but you’re right
>>
File: 311444707646165.mp4 (3.83 MB, 672x768)
3.83 MB
3.83 MB MP4
>>109690134
He has to be stopped.
>>109690144
Yes, but I didn't prompt for dialogue so it's just kinda gibberish. >>>/wsg/6224003
>>
File: Test 00039.mp4 (3.84 MB, 768x1376)
3.84 MB
3.84 MB MP4
>>109688581
Now I need to do a Hamas and Rusich girl to remain impartial in all conflicts.
>>
>>109690124
noonooo anon you need to use it the correct way. there's so much schizo and bullshit from using lodes shit that i just dont bother anymore. most is undocumented and if you arent reading the discord chat daily youll have no clue how to actually make use of it

this is why his shit will never take off
>>
Is it dumb to use H3 for image restoration? Can Krea do it? Can Krea do NSFW?

Anyway so I tried first frame then described an image restoration and that worked to a degree, but what really worked was placing the potato image as last frame slot then describe a pristine image and the process that ruined it. It is a bit of lottery though.

NSFW Example here, I described that the image was cropped and so it gave me frame expansion, but with black bars this time: https://files.catbox.moe/bmizrx.png

There's no nvidia or VSR or upscaling on that. I've also done images that look worse. It seems like H3 can punch through anything.

Prompt, I don't know if an image description is necessary:

Ultra High-resolution 8k 16k scan of the original photograph film negative, digitally mastered in 10-bit 12-bit color and HDR with perfectly preserved highlight detail and shadow detail and smooth gradients, flawless and perfectly preserved texture with incredible micro details like eyelashes and peach fuzz and skin pores, perfectly color balanced and white balanced with accurate skin tones, and perfectly tonemapped with realistic smooth contrast and shades.

A Still image freezeframe of... [Description of the Image]. The scene is frozen completely still and completely motionless, time is stopped.

Then the image is printed on paper with a grainy dot pattern, then that image on paper is aged with shifted and faded colors, then that aged image on paper is photographed with a blurry noisy low quality digital camera with motion blur and inaccurate tinted color and inaccurate contrast with crushed blacks and blown out highlights and blown out color, then that photographed image is downsampled and resolution is greatly reduced and heavily compressed with heavy pixelation and heavy macroblocking and heavy jpeg artifacts and heavy ringing artifacts and low colordepth, then that image is cropped and watermarked, extreme image quality degradation and generational loss.
>>
>>109690061
>you're hitting the nail on the head
Good to know. And no worries, keep your secrets if even the artist's name/art style/period is too touchy. Don't really care about the rest though, just wanted to test my intuition.
>>
>>109690074
Good to see you're still having fun with these.
>>
>>109690239
>most is undocumented and if you arent reading the discord chat daily youll have no clue how to actually make use of it
that explains a lot. v0.2 worked fine so i just assumed he finally figured it out but v0.3 is losing me so far. i'm going to test some other samplers and settings and possibly find that fucking discord you're talking about if i need to.
>>
>>109690086
So? General (merited) disdain for jeets aside, how does them monetizing something you weren't going to hurt you?

>>109690145
Again, not interested so much in the prompt as this one aspect.
>>
>>109690281
I understand where you're coming from anon, I don't think you have any bad intentions
>>
>>109690247
I mean, if it works it works. It's just that you won't be able to gen very high resolutions directly.
>>
File: 751314567473318.mp4 (3.74 MB, 672x928)
3.74 MB
3.74 MB MP4
>>
>>109690074
>>
>>109690183
>It’s not ok anon but you’re right
It is.

I don't have any good answers or great insights but you're okay the way you are even if it doesn't feel that way. I'm not quite sure who decided that this one form of body dysmorphia warranted celebrating mutilation as its "cure" (and fuck whoever it was) but as we know from other forms, indulging it is not the only thing that assuages it. Again, not an expert but perhaps mental exercises that have you intentionally listing and exploring the benefits of having the body you do have could help.
>>
>>109690307
Good one
>>
>>109690202
Marichka, my beloved.
>>
1500 pic collage of training model from scratch.

Pretty happy with what i can pull off with a custom transformer.

It learned the form of people. It did not learn the soul of people though.
>>
>>109690307
>EFF ME
Well, milady, I mean ... if you insist ...
>>
>>109690239
well yeah, that's the purpose of experimental loras. its not for public use.
>>
do you guys write your h3 r2v prompts manually or get an llm to do it?
>>
>>109690369
LLM, though I always double check what it has written and correct where necessary.
>>
>>109690369
i use an llm to lay the groundwork then modify it from there.
>>
>>109690247
>Can Krea do it? Can Krea do NSFW?
flux 2 klein 9b img2img. prompt can be as simple as "clean up this photo" but i sometimes mess around with adding more instructions depending on the output, sometimes they work out better. it's fun but my expectations are highly tempered.
>>
>>109690349
neat
>>
Thoughts?

https://civitai.red/models/2889902/painting-style-anima?modelVersionId=3267128
>>
No one cares about babbies first [model] LoRA
>>
File: test_00078_.png (1.02 MB, 1536x1536)
1.02 MB PNG
How would you guys write the H3 prompt to account for a reference sheet like this? How do you ensure it applies information from every view to subject x?
>>
>>109690124
>>109690239
there's a different version of kroma that someone posted in the last thread that's better than the one you're complaining about.

but yeah, you have to join his discord if you want to keep up with it unfortunately
>>
>>109690403
>Anima
anon it's august, we all use h3
>>
>>109690247
Edit: you want to describe the image or drop "eyelashes and peach fuzz and skin pores" because you'll get a picture of an eye on a random roll.
>>
>>109690413
<Subject 1> is the girl in <Picture 1>.

I'd try just that at first, if that doesn't work maybe describe her appearence succinctly.
>>
>>109690413
The regular prompt should work, but if it's not holding the character or hallucinating multiple ntitties try:
<Subject 1> is a girl in <Picture 1>, she is wearing a black schoolgirl uniform, she has black hair and gray eyes. She is depicted in a closeup frontview on the left in <Picture 1> (if you had a closeup view) and then from a frontview, sideview, and backview going from left to right on <Picture 1>.
>>
>>109690424
You should really try to describe her, otherwise the model might get details wrong.
>>
File: i_00464_.png (1.34 MB, 1280x720)
1.34 MB PNG
>>109690202
why tf are her tits so small?
>>
>>109690413
subject_definitions:
<Subject 1> is the same anime girl shown across the multi-view character reference sheet in <Picture 1>. The front, profile, and rear depictions are complementary views of one subject, not separate people. Combine the visible information across these views to preserve her facial identity, wavy dark hair, body proportions, sailor-style school uniform, and view-specific hair and clothing details.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the identity and appearance established across all views in <Picture 1>, including view-specific hair, silhouette, and wardrobe details whenever those angles are visible.
>>
File: collagetest.png (1.64 MB, 1877x1364)
1.64 MB PNG
>>109690413
Should just werk by declaring the image is the subject, I've done this single image collage test myself. Just know you are sacrificing resolution by dividing the image into multiple views or you need to use the slow max res mode.

I think you only need to describe appearance if something isn't sticking or there are multiple potential subjects in the image and you want to help pick the subject out.
>>
>>109690413
You only need 1 close up face shot, full body front and back. Side view is not necessary if the design is generic.
Tell H3 that <picture 1> is a character sheet of the same character, and she is <subject 1>. Fully_preseve <subject 1> facial features from the close-up shot of <picture 1>
>>
>>109690202
>>109690441
nazi shit isn't cool.
>>>/pol/
>>
Also anons, you can name your subjects like <chad> and <stacy> instead of <subject 1> and <subject 2>. It makes it a lot easier to type the prompt.

You can also define subjects without references. They are also good for long descriptions of characters and environments that would normally break the prompt if it was shoved into the main body and the action or camera starts focusing on your descriptions. Just make sure to use subject_definitions: and whatever the main body header is called, it's different for fl2v and ref2v.

So you can do <stacy> is doing blah blah blah in <fieldhouse> with <chad> <brad> <aaron> <nick> and they are all autistically described, but the main action and the camera is controlled by a single simple sentence.
>>
>>109690482
Minimax's own SKILL.md says you should have the following information

Each character card should be a 16:9 production reference sheet when possible. Unlike final rendered video, character cards should include clear readable labels so downstream generation can bind the correct person and props:

Character name label in English and/or the project language
Role label, such as protagonist, grandma, thief, sidekick, pressure character
Main 3/4 view
Front / side / back views
Expressions
Material / costume / prop details
Important prop labels, such as handbag, wallet, skateboard, apple basket, scarf, shoes, glasses
Identity lock repeated in the prompt
A short visual-ID note listing age range, body type, hairstyle, outfit colors, signature props, and do-not-change traits


https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/3d-animation-short-generator/SKILL.md#step-3-character-cards
>>
File: Return_00249_.jpg (1.53 MB, 2728x1536)
1.53 MB JPG
>>
for the anons who made their own krea 2 + kroma 0.3 merge, what weight ratio do you recommend for these two model to retain krea 2 knowledge and benefit from kroma nsfw understanding?
>>
>>109690441
Because I prefer it that way, and that's also how most depictions show her.

>>109690489
I'm not endorsing anything, just wanted to gen some evil looking women.
>>
File: Return_00254_.jpg (1.49 MB, 2368x1776)
1.49 MB JPG
>>109690516
I think most anons are waiting for them to sort that out, I haven't seen any gens from the merge
>>
File: Lookin' Sharpe.png (1.89 MB, 1152x640)
1.89 MB PNG
>>109686128
>Morning, anon.
>Good morning, saar.
>Where is anon!?
>I am Dikshit. Brahmin, CEO of Scams Incorporated. Until 3 days ago, dalit Pajeet. Outside is my partner in crime, a troll you reported through the threads.
>>
>>109690500
You don't need all of those.
Expressions are good, though.
>>
>>109690489
>nazi shit isn't cool.
Says who?

>>109690501
>>109690532
I wanna fuck that fox lady right in the pussy, as a legend once put it.
>>
Another Uncanny Valley Sunday.
>>
>>109690544
Either that or
>I am Dikshit. Brahmin, janitor of /g/. Until 3 days ago, dalit lurker Pajeet. Outside is my /ldg/ OP, a troll you reported through the threads.
>>
Anyone know a good eye adetailer for krea2?
>>
>>
File: i_00568_.png (3.83 MB, 1536x2048)
3.83 MB PNG
>>109690489
>nazi shit isn't cool.
why?
>>
>>109690629
hot. this makes me wanna JORK it
>>
>>109690629
because there are jewish people that browse this board
>>
https://github.com/whp199/GemmaPrompt
https://github.com/InsertSpice/MiniConstruct

so which ldg vibememe is better?
>>
>>109690629
I'm surprised it got the Luger kinda right.
>>
>>109690532
damn it, ok
>>
>>109690625
Pineapples on pizza is an overrated debate.
>>
>>109690644
there's like 20 of these. everyone here is probably using their own methods
>>
>>109690668
most of them are custom jeetnodes though.
>>
>>109690687
what's even a jeetnode, do you mean they're broken?
>>
Any monkey that can run qwen 27b at Q5 or higher can make these nodes
>>
>>109690124
its not user friendly atm. Follow his discord to know how to use it properly
>>
>>109690644
MiniC doesn't have an unload option for retards like me that keep forgetting to do it when finishing the prompting phase.
>>
File: 263210285345524.mp4 (3.62 MB, 672x928)
3.62 MB
3.62 MB MP4
>>
>>109690709
It means poorly documented vague shitnodes that create dependency hell and don't fucking work half the time
>>
File: 1098080218031611.mp4 (3.63 MB, 1056x608)
3.63 MB
3.63 MB MP4
>>109690413
This one is using
>subject_definitions:
><Subject 1> is the girl in <Picture 1>, a Japanese student with short dark hair, wearing a navy blue sailor-style school uniform with a red ribbon and grey socks.
><Picture 1> is the reference sheet showing the character's design from multiple angles.
>>
>Follow his discord to know how to use it properly
sigh...
>>
File: 1608429165255.png (45 KB, 789x750)
45 KB PNG
>>109690722
well, i guess that will be in the next upstream commit.
i usually just eject in unsloth
>>
>>109690784
You have trained yourself well but these other tools have made me lazy since they offer this option and I've gotten used to it. I'm sorry.
>>
>>109690644
I just use the native "Generate Text" node.
>>
File: 482491070648133.mp4 (3.67 MB, 1056x608)
3.67 MB
3.67 MB MP4
>>109690413
>>109690429
>>109690779
This one is just
>subject_definitions:
><Subject 1> is the girl in <Picture 1>.
>>
>>109690811
looks like you changed the detailed description
>>
>>109690821
No, the rest of the prompt is exactly the same, different seed though.
>>
File: screenshot.1788122962.jpg (21 KB, 366x134)
21 KB JPG
>>109690792
((MY)) prompt enhancer already had an unload button. It also has a custom node for Comfy that automatically unloads any ollama model if a workflow is ran, so even if you forget to click the button in the prompt enhancer, it does it for you.

This works BOTH WAYS TOO! If Comfy is currently using 100% of your vram and Comfy has models in there, using the prompt enhancer will slow your system to a crawl. The prompt enhancer checks vram usage and warn you if usage is about 20%.

((MY)) prompt enhancer has already handled these edge cases because im so based.
>>
File: Test 00040.mp4 (3.75 MB, 1200x670)
3.75 MB
3.75 MB MP4
>>
>>109690911
we evolved from
>1girl, standing
to
>1girl, walking
and all it takes is a 5090 with minimax raping it for 10 mins
>>
File: Return_00278_.jpg (1.3 MB, 2368x1776)
1.3 MB JPG
>>109690930
>10 minutes
Not at that resolution I can get it under 5
>>
>>109690947
it's a bad thing because it enables your spamming. I hope you get a virus from comfy
>>
File: 607653128984800.mp4 (3.82 MB, 960x544)
3.82 MB
3.82 MB MP4
>>109690930
>rape
more like lovey-dovey sex
>>
File: zel4.webm (1.95 MB, 928x672)
1.95 MB
1.95 MB WEBM
>>
>>109690958
sound? at least put the jew song in it
>>
>>109690963
Forehead like a billboard, she still cute though.
>>
File: concept778.mp4 (3.86 MB, 1376x768)
3.86 MB
3.86 MB MP4
>>109690953
You can keep crying while we thrive here
>>
>>109690930
Neat, innit?

Obviously didn't take 10 minutes. Generated at 1MP and 8 steps in about 7 minutes. Had to hit it with a slight downgrade in resolution and CRF 26 to fit here, though.
>>
>>109690978
What's the point of using H3 if it looks like ltx?
Also you've been rate limited, when I try to open any of your reposted images or videos they have issues loading unlike the others in thread, maybe you should take that as a sign
>>
i don't want to have to pay for nudify services and risk a knock on the door from police
will a llm fix my issue
>>
File: 3061072000.mp4 (3.73 MB, 800x768)
3.73 MB
3.73 MB MP4
>>
File: output_smal.mp4 (3.92 MB, 2048x1130)
3.92 MB
3.92 MB MP4
>>109691003
Meds lil buddy
>>
>>109691008
It's quite ironic that you of all people tell people to take meds :^)
>>
Does it matter for h3 to give it a character sheet with front, back etc angles instead of just a standing front character in terms of being able to make an accurate video, or not really?
>>
>>109690643
They'll get over it
>>
>>109691042
its more likely that a law gets passed to prevent you from doing it
>>
>>109691038
If the character turns around, then you're going to want to ensure there is consistency with your vision.
I can also tell you for 100% a face fidelity reference is very important for capturing likeness, especially with life action characters.
>>
>>109691038
Doesn't matter if you just want to do a single gen with that character. Minimax will just make shit up.

Only matters, if you want to have a consistent character accross multiple gens, because Minimax will always make up slightly different stuff.
>>
>>109690989
7 minutes is still a lot of time just to see an animoo girl walking or cooking like your other gen. but whatever floats your boat.
>>
>>109691048
>>109691054
I see, thanks anon, guess I'll try to do that then, since the model understands it.
>>
>>109691044
That sounds like terrible PR for jews and a great example for antisemites to exploit.
>>
>>109691061
Are you bitter that your rig can't do it at all?
>>
File: tree.jpg (3.42 MB, 3398x4518)
3.42 MB JPG
>>109691038
>>109691054
what he said
>>
>>109690349
And here is the final output with Euler Sampling implemented, SD 1.5 duct taped at home.

Im decently impressed with the result. This model weights were trained from scratch while i slept the past week on my 3090.

It knows a decent amount and i think it captured <1girl, looking at viewer, animal ears, smile, breasts, dress, long hair, gloves> really well with only 1 million input images.

Final parameters came out to be 733.6M which is very close to standard Stable Diffusion 1.5 which i think is 800ish.

My biggest take away is that more images != more better. I think if i ironed out a vision model to identify high quality i could get the same results with 100k images then lora train specifics on small datasets to teach it who Midna is.

I think anons thinking you need 100+ million images to train something is chinese propaganda so people do not try making their own models.

if you gave me a Data Center and ill create a model only trained on opensource anime art then people wont be able to complain about stolen art anymore.
>>
>>109691070
i dont do video gen, its a waste of time. my rig is probably better than yours unless you have a blackwell 6000.
>>
>>109691061
Maybe, but I enjoy those comfy gens. Can just let them run in the background.
>>
File: file.png (128 KB, 1701x807)
128 KB PNG
returning to after a while, decided to try anima and realized that forge classic doesn't recognize it and that it's no longer on active development, so I Installed neo forge/ forge classic.

when I finally got it running I got this error on the console, I was able to run auto1111, comfy and 2 versions of forge before normally, does anyone know how to fix this?

i have 6gb vram and 16gb ram.
>>
>>109691090
Then post it instead of being a crying bitch
>>
>>109691105
post what?
>>
been out of the loop for a lil bit. can anons point me to which lora is best to jailbreak krea 2?
>>
>>109691088
Pretty neat.

People complaining about stolen art will find some other thing about AI they can complain about. No reason to acquiesce to them.
>>
File: 360460506515611.mp4 (3.7 MB, 672x928)
3.7 MB
3.7 MB MP4
>>
>>109691133
textfusion refusial reduction at lora strength 1
>>
>>109691133
>>109691142
some anons here said it was thing one :
https://civitai.red/models/2746817/krea2-filter-bypass-fedor
>>
>>109691152
*this
>>
>>109691133
do you even need a lora for that with finetunes?
>>
files.catbox.moe/7j1wmr.jpg
>>
>>109691088
What a waste of electricity.
>>
>1girl, standing
>watermark
embarrassing
>>
>>109691152
>>109691133
just tested this for fun against textfusion refusal reduction and it didn't work and left the image censored fyi
>>
>>109691178
nice, but she should stop smoking
>>
>>109691196
too many nogens
>>
>>109691206
they're herbal cigarettes like the ones they use on film sets. don't worry muchacho!
>>
>garbage gens everywhere
>>
fucking simpletons, you could not fathom the intricacies of my 1girls...who are standing.
>>
>>109691205
ok, I was convinced by the sales pitch of their description too
>>
>>109691178
>poltpao.koim
krea 2?
>>
>>109691226
>>109691220
We have a lot of diverse gens what do you want to see exactly seether kun?
>>
>>109691220
check out >>109680031 for the good gens
>>
>>109691080
>>109691220
what the fuck did you just say about my 1tree gen
>>109691229
Hey take whatever you want, but fedor filter bypass is outdated and doesn't even work, it's a 32kb download just test it for yourself
>>
>>109691152
Krea needs filter to bypass? I've had more issue with klein models than Krea.
>>
>>109690521
Swastikas aren't evil, they're a cross.
>>
>>109691235
k2+zit
>>
>>109691242
i havent used the bypass and it works fine. they must be using the base model.
>>
>>109690779
Does she get isekai'd?
>>
>>109691237
gen a spooky skinwalker with antlers peeking from a tree, he's staring at a 1girl, standing
>>
>>109691038
with a single front facing character image it wants to use it as an exact frame somewhere. with ref sheet it understands better that its just a reference
>>
>>109691268
Oh that's cool anon
How about you do it instead
>>
>>109691220
>entitled no-gens everywhere
>>
>>109691252
krea 2 raw then zit?
never tried this combo
>>
>>109691277
i already did, but you asked me what i wanted to see, so i told you.
>>
/sdg/ is actually a fucking wasteland holy
>>
>>109691280
>anon thinks someone no posting their gens means they're a no gen
this thread should primarily be for the latest local image/video model discussion. any gen posted here is for testing purposes.
>>
>>109690025
>>
>trying to bring back "nogen"
Sorry anon that doesn't work here
>>
>>109691252
you used two models and 1girl standing in a convenience store holding a jack daniel's is all you got? KEK
>>
File: Return_00304_.jpg (1.12 MB, 1776x2368)
1.12 MB JPG
>>109691285
Yeah it's mostly the same guy pretending to be other anons and 4 other dents that post from time to time
>>109691296
It does work when they bitch and moan about what anons post.
>>
File: file.webm (2.88 MB, 640x640)
2.88 MB
2.88 MB WEBM
Mooo!
https://files.catbox.moe/7x7u9l.mp4
>>
>>109691281
its very good. zit gives k2 detail

files.catbox.moe/xlm4xw.jpg

>>109691297
I dont give a shit about u nogen

>>109691305
damn nice!
>>
>>109691305
can i get a catbox for that? looks nice, is that the muscular oni from touhou?
>>
>>109691141
>five frames later
>who?
>>
>>109691335
I can't give out box but yes it's her, I'm shit testing the model and working out ways to circumvent characters from altering the style so far it's been working but I'm going down common characters
>>109691321
Krea2 with that anon booru lora really allows me to go back to my bread and butter, doesn't seem to damage the styles I like
>>
>>109691286
Says who?
>>
>>109691353
which boora tag lora?
>>
>>109691353
>Krea2 with that anon booru lora
?
also does it understand artists and does it absolutely needs tag style writing?
>>
>>109691353
>I can't give out box but yes it's her,
alright, can you at least tell me if its krea2 and the prompt you are using to get that style?

looks like Bouguereau
>>
File: Myspace526.mp4 (722 KB, 704x288)
722 KB
722 KB MP4
Behold, this is what boomers will use ai for in 10 years
https://litter.catbox.moe/u0lrxh.mp4
>>
>>109691372
all vtuber/streamer fanatics should be publicly shamed and made fun off until their shame consumes them

Monster Hunter Wilds is a video game.
>>
>>109691370
No it just gives concepts
>>109691363
Quarterturn lora he posted it in the thread check hugging face or the archive
>>109691372
Sorry legit there are people that sit here 24/7 just seething at me trying to screw me over.
>>
>>109691311
Cute (and sexy)! And what a positive outlook on life this buxom young bovine has, always looking at the bright side!
>>
>>109691321
Did someone yeet a glass of marmalade at her?
>>
>>109691381
>Sorry legit there are people that sit here 24/7 just seething at me trying to screw me over.
i just asked for you the style prompt you used since you're not willing to provide a catbox. i don't know what the fuck you're talking about or care but sure, hoard your prompt then. some of you people here are very mentally ill. i'm gonna go jack off to the oni 1girl.
>>
>>109691404
Enjoy!
>>
>>109691410
i will, fuck you.
>>
>>109691404
>i'm gonna go jack off to the oni 1girl.
Do NOT download this collection of lewds then:
>https://files.catbox.moe/t7eab5.zip
Seriously, don't. This isn't reverse psychology. You have been warned. This link is to be avoided like the plague.
>>
files.catbox.moe/zd8z0a.jpg
>>
>>109691404
I'm just using the atlyss style + the mega man psX loras for SDXL with various furry models from cvitai + the style guidance prompt: (((masterpiece, very aesthetic, dynamic lighting, muted color palette, Atlyssstyle, retro, lowpoly, 64 bit, N64, PSX, PS1, 3D aesthetic))). (((lowpoly 3D n64 PSX PS1 pixelart themed feral hybrid animal bovine moo cow quadrupedal quadruped))).
>>
https://files.catbox.moe/k7ohkm.mp4
>>
>>109691494
Tasteful.

>>109691513
jej
>>
https://github.com/Comfy-Org/ComfyUI_frontend/issues/4195
I thought I was going insane and missing something but this is actually not a feature in Comfy? What the fuck are these retards doing?
>>
File: restored3_censored.png (3.73 MB, 2047x1344)
3.73 MB PNG
>>109690247
Just got this result

NSFW: https://files.catbox.moe/fzhj45.png

Not only is everything cleaned, upscaled, colorized, but the highlights that were blown out in the window are completely restored. This means you can properly HDR all content with bright detailed highlights.
>>
>>109691152
This one can destroy prompt adherence too much.
>>
>>109691550
Kind of amazing how it can look so good and natural, good job anon.
>>
>>109690247
I've used H3 reference model to successfully recreate highly accurate turnaround vids of multiple people from decades old small and noisy photos/footage, with different hairstyles and outfits. So no, it's not dumb. I haven't tried Krea.
>>
>>109691635
>boomers died just before this technology
>>
>>109691658
they also died before they got to see the collapse of society. lucky them.
>>
I like that patchouli, I saved her. I hope you guys don't mind.
>>
>>109691550
Actually sick, awesome work anon. Seeing cool ways to explore what H3 can do is why I'm here.
now if you'll excuse me it's back to 1girl tentacle fuckmachine gens
>>
>>109691550
thats pretty good, whats was your workflow?
>>
>>109691100
I'm having a similar problem with GTX1050
Anyone knows?
>>
>>109691658
>>109691690
Yes. The future and technology are double-edged swords. Enjoy them while we still can.
>>
>>109691783
just the default i2v template with the prompt posted earlier. I didn't even describe the image because I was doing them in the background. Just remember that the potato image goes in last frame not first frame
>>
>>109689792
How can I make realistic face swaps with generated characters? I want to deepfake a character I created but make it look as real as possible
>>
>>109691940
We don't do that here
>>
>>109691550
that's really clever.
>>
>>109690318
Thanks for your kind words i guess it’s something I’ll have to deal with since im already quite old to do something retarded.
>>
Wanna swap faces, anyone?
>>
>>109691550
We're used to decadent and depraved 4k pornography on download. Think about how this dame was photographed when color or hi-def media were unicorns and you could only see it as a G rated picture at a high end theater. On the left is how you would see bitches in magazines. Imagine that blurshit is your baseline and how you imagine hot women, now imagine what its like to be in the room with these missile titties hitting your retinas. And you can't fuck her because this is before the pill and you're not rich and you're just the photographer. Moment of silence for this man.
>>
File: face-off-lead.jpg (26 KB, 1200x680)
26 KB JPG
>>109691984
>>
File: Return_00405_.jpg (1.18 MB, 1776x2368)
1.18 MB JPG
>>
>>109689970
It's not clear what this tool does? Are there examples of how to use this?
>>
Is the default workflow for h3 still good? Any turbo loras or extra acceleration for both models that are any good?
>>
>>109692006
Use sageattention for speed boost. Turbo loras destroy prompt adherence and output quality.
>>
>>109691088
Somebody get this man a data center!
>>
>>109692028
I mean how bad is the trade off? I'm currently at 5~ minutes for a 10 second gen
>>
>>109692036
you can just...test it to find out
>>
>>109692036
I used to wait 1 hour for a 5 second gen.
>>
Is there a realistic lora that is "unrealistic"? I mean like the instagram/tiktok filter stuff. The chinese use them, but I don't wanna make account on modelscope.
>>
>>109692006
>>109692028
I'd recommend the built-in CK attention. It's a bit faster for me.

Turbo loras can be fine, depending on what you want to generate. If it's not all that complicated or has fast movement, just slap a turbo lora on it.

I think the main timesaver is having a preview with taeh3, because you can abort gens, that are obviously not what you wanted or don't follow the prompt.
>>
>>109690811
I would love to see a blood splatter during that slo-mo sequence and also once that toast hits the ground like a bucket of blood splatter over it.
>>
>>109692028
>>109692036
You can also just use comfy kitchen, I think its always a bit faster and you don't have to install it.

Trade off is unanswerable because it depends on what you are trying to make. If something is short and simple the result isn't bad. Just turn speed and cope nodes off when you aren't getting results.

Also for 10 seconds you can use sparse attention, but it does fuck with adherence if your prompt isn't tightly written, and complex motion is fucked, and if there's major changes from the beginning to end of the clip I wouldn't use it. But it will put out a 15 second clip pretty fast, like a 2x+ speedup on long gens.
>>
File: Return_00418_.jpg (1.24 MB, 1776x2368)
1.24 MB JPG
>>
*vomits*
>>
>>109691286
>this thread should primarily be for the latest local image/video model discussion. any gen posted here is for testing purposes.
nuh uh, its for mentally ill homosexuals who need to post their 1girls for other men to see and avatarfag
>>
I get into genning off an on, I often take breaks for months then come back if I'm paying attention around the time a new model hits. So I have huge gaps in my knowledge of developments here.

Anyway a while ago, at least as far back as last year, maybe more, I remember /ldg/ having a culture that disparaged the posting of 1girls.

Anyway so my next question is when did the fag leave lmao
>>
>>109692168
nobody left. same fags still here. they're probably never going away.
>>
>>109689792
Why is she white? Patchy is Japanese.
>>
>>
>>109691920
>>
In the near future when video models are good enough to generate several minutes and understand incredibly long and intricate prompts, a new type of competition/pastime will develop:one shot prompting to see who makes the best short film based on a given premise. You would need the skills to preemptively predict the type of mistakes the model could make, know where it needs extra handholding, how to preserve continuity only via prompting, be able to envision how your prompt would look like without aid, and understand all the quirks of the model. You would need to write out everything in one long prompt, including specifying the shot, framing, composition, lighting, dialogue, costumes, background, setting, screenplay, etc
There would be a panel of human judges to see who has the best one shot short film.
>>
>went through like 30 prompt variations trying to get the guy to just start pleading directly after she orgasm denials him
this model sucks man, timestamping and shot separation dont work at all, the guy always has precognition.
>>
>>109692470
heh, you're getting your own orgasm denied because the guy who's cucking you is not playing his part correctly, kek.
>>
>>109692464
mate you know some retard will just be able to say "make a new episode of seinfeld" and the AI will just be able to do it with an LLM and it will be style and convention correct in every way and you can't tell an AI wrote the jokes. Illustration just completely died as a market valued skill despite being around since cave paintings. How long do you think proompting will last?
>>
>>109692470
Well if you're up to it, you can post the prompt. My guess is your prompt is stamped correctly but the prompt body is too long with detail descriptions.
>>
>>109692524
I wasn’t trying to prop up prompt engineering as some kind of new skill. Just proposing a new sport. Humans can play chess even with superhuman engines. In fact you can imagine a model fine tuned for the purposes of the competition so that anything underspecified would end up looking bland or shit
>>
>>109692566
>walk in holodeck
>computer, make yourself retarded
>now do exactly as I say

holy shit did you hear what chad pulled off in the holodeck yesterday
>>
>>109692579
Who the fuck is chad? Stop making up dumb shit
>>
>>109690247
How many steps are you using?
>>
>>109692585
cucks can only think of chad doing things, please understand that most of this general is brown.
>>
>>109692524
You have to be fucking braindead if your honeymoon phase with those seinfeld slop shorts lasted more than two minutes.
>>
File: 1779705771929.png (1.89 MB, 1416x1416)
1.89 MB PNG
>>
>>109692524
I doubt frontier LLMs (which are corpo-centric) will ever be good at humor.
For things to be funny you need some degree of political incorrectness which only uncensored models can do and most uncensored/"abliterated" models are retarded
>>
>>109690413
white background might affect lighting should use grey instead
read on reddit
>>
What’s stopping me from using atlas maps if we can use character sheets? Why say it supports a max of nine image refs
>>
>>109690644
> MiniConstruct
> looks much better
> dedicated h3 tool
> works with other llms
> not abandonware
> author knows how to gen
i think it's obvious
>>
>>109692585
Exactly, no one will know who you are for "prompting skills" speaking of making up dumb shit
>>
>>109692589
20. Sometimes it doesn't go well
>>
>>109692672
>literally said there is no skill in prompting
Dumb esl faggot.
>>
>using comfyui templates
>try using qwen3vi to identify lewd images
>just keeps saying "sorry i cannot fulfill that request"
anyone got a jailbreak prompt?
>>
>>109692596
>honeymoon phase with those seinfeld slop shorts lasted more than two minutes.

You know that basically makes it perfect to insert in a satirical sentence like the one you just replied to
>>
for h3 ref2va whats my best course for going past 15 seconds with consistency?
i can obviously do 15 seconds, then throw the output into it and extend it, but that seems like a really inefficient way to do it. would it be better to try do a ref2va after the first gen then treat the final frame like the first frame and just reuse the same refs? its not clear what the best way to go about this is
>>
>>109692694
>I was pretending to be a retard all along
That’s cool kid
>>
Is Mitchell better than Lanczos?
>>
>>109692606
Grok is less constrained and you can deliberately ask it to be edgy. Besides there's a lot ways to be funny that aren't politically incorrect or adult content. Humor is just apparently difficult period. LLMs are also bad at charisma even though they can imitate voice actor delivered human speech with audio pretty well.
>>
>>109692470
>zoomie can’t read a fucking guide
>>
File: ComfyUI_16667_.png (1.56 MB, 1024x1536)
1.56 MB PNG
>>
>>109692684
>no skill but here's my idea for a sport

Look dude, I wasn't even trying to make fun of you or your ideas but now you're just being a disingenuous faggot about something nobody needs to argue about. It's fine go do something else.
>>
>>109692706
Yeah. Obviously if I could 1 shot prompt a TV show it would be friends
>>
File: x_80poea.png (568 KB, 1024x1024)
568 KB PNG
>>
>>109692737
Maybe you’re fucking autistic so I have to spell it out for you. My point is that in a world where AI can generate anything, the only “skill” one has is natural language. So a pastime might develop where people will compete in that. Obviously the unspoken part is that an environment and ruleset have to be set up so that humans are competing with each other. You went off on some inane tangent about art that’s as irrelevant as saying computers can beat humans at chess, because your reading comprehension is so poor you thought I was saying proompting is some hot new skill.
>>
>>109692764
>maybe you're autistic so let me type all this autistic shit again
>>
>>109692680
Sorry, I am retarded, I meant duration (seconds). How many seconds are you using?
>>
>>109692799
4 seconds because its the minimum spec, I don't think you should be able to do less.
>>
File: Return_00433_.jpg (1.32 MB, 1776x2368)
1.32 MB JPG
>>
File: ComfyUI_00019_.png (2.01 MB, 1536x1536)
2.01 MB PNG
>>
>>109692811
your brain is buck broken
>>
>>109690544
one does not simply march into Anatolia
>>
File: Return_00439_.jpg (1.29 MB, 1776x2368)
1.29 MB JPG
>>109692853
>>
>>109692724
>Besides there's a lot ways to be funny that aren't politically incorrect or adult content
Doubt it.
Even in kids shows there used to be jokes about certain characters being fat and/or ugly and even those things are considered to be politically incorrect these days.
>>
so I finally set up a local model on my modest PC (32gb ddr4, 5070ti) and after playing around with the popular censored online models last year it is currently blowing my mind what you can do with this shit locally
just straight up creating gooner shit in minutes, nothing's censored
it's crossed my why I shouldn't just try to go full pajeet mode and set up a side business with this to scam coomers (free shit on twitter etc that links to a patreon that does commissions)
>buzz lightyear on shelf.gif
I'm tired of my 9 to 5, I get paid peanuts here in Germany despite having a masters in comp sci
>>
thanks for reading mein blog btw, sieg
>>
>>109692902
ja
>>
File: ComfyUI_00030_.jpg (2.48 MB, 3840x3840)
2.48 MB JPG
>>109692895
never goon
>>
>>109692895
>everything is going great, coomers are lining up
>the coomers are just chat bots
>>
>>109692853
no reason in replying to him. he thinks everyone is the one guy he hates.
>>
>>109692895
dont forget to melden your coomer einkommen to your local Finanzamt so they can calculate the Einkommensteuer you owe them.
>>
>>109692650
Character sheets don't work well for facial likeness.
>>
>>109692464
this can already happen now with images which are much smaller and people dont want to waste time doing it, let alone video
>>
File: ComfyUI_00050_.jpg (1.45 MB, 2080x3040)
1.45 MB JPG
>>
Please don't post gayness. It's against my religion.
>>
>There isn't a girl out there that's genning a 1boy that's just like you
>>
>>109693040
damn thats crazy, she should make better gens
>>
Has anyone else starting counting cuts in TV Shows since H3 got released? I counted 8 seconds as a longest shot in an episode earlier and it made me think.
>>
>>109692736
You can't post this kinda fire and not share the prompt, anon
>>
>>109693040
Yeah no.

women care about looks more than men care about sex.

The incel truth: when you look in the mirror or ask your mom you get really unreliable face rankings. If women don't basically pester you, you're not good looking. No amount of Jesus makes a woman anything different than "of the flesh". All women are such, most men are too.
>>
>>109693040
speak for yourself
>>
File: ComfyUI_Krea_2_00427_.png (2.12 MB, 1152x1152)
2.12 MB PNG
>>109689792
ACE-Step XL Yousei Teikoku LoRA
https://files.catbox.moe/inbl3t.flac
https://files.catbox.moe/o5gsag.flac
https://files.catbox.moe/m42nvw.flac
https://files.catbox.moe/zczuto.flac

Using the VAE from
https://huggingface.co/mdmachine/ACEStep-XL-Regrind-V1
Results seem to be way better on the 0.3 merge model.

ACE-Step XL still surprises me with what it can do, a step above Minimax Music for now since it can do LoRAs.
>>
>>109693063
It's just not true. Women aren't like that. Women like to produce results through patterns of putting on, but there's no conception of anything like you could have. They'll be as much the same wherever irrespective of the you that you are, you have nothing to do with their patterns. They simply engage in them in accordance to whatever their desires are, and are an existence utterly apart from your own entirely.
>>
>>109693073
not to shit in your party bowl or anything but it still sounds like shit from an ass. no offense.
>>
>>109690247
If it helps anyone I kept working on my prompt. Note I don't actually know if it works better, I just added or changed terms when I had a thought. Prompt in next post. Things that I think might help:

1) Putting the potato image transition in [Shot 2]

2) Showing sound prompts are "N/A.", I know sound can change video results so try to turn it off.

3) Adding a black and white sentence, you should add or remove this depending on if the image is black and white, as well as the description of "color photograph".

4) If you have time I would describe the image after the high fidelity image boilerplate. Should help with recognition for details, colorization, frame expansion.

Also all these sentences are on an add/remove basis and ideally you'd want to take out sentences that don't describe the potato image, this is just everything for catchall. Also I have no idea if these sentences or terms are understood, arguable any time it reads a prompt it doesn't understand, the prompt is less effective. Length =/= better either.
>>
>>109693073
>https://files.catbox.moe/inbl3t.flac
which settings do you tend to use?

I found out if I put the steps in the little custom box in ace step cpp I can do 500 steps. But rn I'm just doing (pasted) simple.

appreciate the new vae.

but which again is the merge? I remember you posting that.

>>109693073
>ACE-Step XL still surprises me with what it can do
I am pretty sure it could be trained to do this:
midi to wav, some melody
load it
cover mode
prompt just has appropriate lyrics
->it sings the melody from the midi

I'm pretty sure this can happen. get it to do it ... infrequently, in repaint, and in cover.

idk maybe cover is a better choice for training? I've never tried to do this, idk if a lora could do it, or if it's more a finetune. Your feedback would be welcome, because

>no model in the world yet does midi + lyrics -> ai music
>>
>>109693102
integrated_multimodal_description:

[Shot 1]

Ultra high-resolution 8k 16k archival scan of the original film negative of large-format color photograph, digitally mastered in 10-bit 12-bit color and HDR with perfectly preserved highlight detail and shadow detail and tonal detail and smooth gradients, flawless and perfectly preserved texture with incredible fine detail and sharp edge clarity, perfectly color balanced and white balanced with accurate skin tones, and perfectly tonemapped with realistic smooth contrast and shades.

A still image freezeframe, the scene is frozen completely still and completely motionless, time is stopped.

[Shot 2]

Then the image is decolorized and made monochrome in black and white,

Then the image is poorly transfered with added filmgrain and dirt and hair and scratches, then that image is printed on paper with lost detail and dithered halftone dots, then that image is damaged and photoaged with shifted faded colors, then that image is photographed crooked with a low-quality digital camera blurry out-of-focus with motion blur and color noise, then that image is poorly digitized with inaccurate tinted color hue rotation and crushed blacks clipped blacks and blown out highlights clipped whites and blown out colors clipped colors, then that image is downsampled nearest-neighbor and made low-resolution with aliased pixelation, then that image is heavily compressed with compression artifacts and heavy macroblocking artifacts and heavy jpeg artifacts and heavy ringing artifacts and heavy sharpening artifacts and low colordepth, then that image is watermarked with a logo, severe image quality degradation and generational loss.

overall_soundscape:

N/A.

non_diegetic_music:

N/A.
>>
>>109693073
also, the song is highly listenable. I have no idea what she's saying, because I don't know any japanese except hentai and kimono.
>>
>>109693073
Instruments are mostly fine, songs are fine, but the voice is still bad, it's like her voice has been sampled from a bad 64kpbs mp3.
It's especially the case with the second song.
>>
>>109693063
who are you genning femanon
>>
>>109693059
its over...
>>
>>109693115
no it isn't.

You don't have golden ears like me. That's not a correct description of the sounds.
>>
>>109693081
oh sorry, I misunderstood your post. I thought it was a "I hate myself because I'm a low value male", but it was actually a "I hate all women because they don't behave how I want"
>>
>>109693126
It was never over. We just need Christian patriarchy. Why do you think everyone married and nobody divorced in the 18th century?
>>
>>109693128
Whatever you say anon, her voice sounds bad anyway.
>>
File: evelyn_west_censored.jpg (754 KB, 2067x1356)
754 KB JPG
>>109693102
>>109693107
Color descriptions seem necessary, I don't know why it didn't understand mommy's nipples are always pink

catbox NSFW: https://files.catbox.moe/9115o8.png
>>
>>109693073
ACE-Step now can finally do proper layered vocals with this LoRA.

>>109693115
Voice in real metal songs tends to not be that crisp, listen to some of training data I used
https://www.youtube.com/watch?v=8Sn0PSAqPn8&list=OLAK5uy_kRToFFGT6ajftoBsXMThIH_dXpGFOkEug&index=4

https://www.youtube.com/watch?v=NS3MPV2yCgo

Listening on HD600s they both sound pretty close. Though I'd agree the biggest weakness is voice right now (not perfect), but these songs are pre-mastered, running them through Matchering 2 should improve that, and I'm assuming feature improvements to ACEStep will also help with that (the improved VAE really removed a shit ton of artifacts those type of complex gens have)

>>109693105
I might create a inference/updated training rentry later, this is using the 0.3 Base/Turbo merged model, DiT-only with LM disabled, CFG of 20, only 50 steps (I do not take more steps), MD Storm 4 sampler (which is best), V10 Regrind VAE I linked, everything else is default
>>
>>109693073
Why not use Minimax Music? it has much better quality and natural sounding vocals.
https://files.catbox.moe/6ndbuy.mp3
>>
>>109693132
I was going to respond, but look, you do you or whatever, good luck (you'll fail and not know why).
>>
>>109693186
>feature improvements to ACEStep will also help with that
future*
>>
>>109693139
You're not Japanese. You're not a musician. You are not relevant, your opinions don't matter, you are nothing utterly.

>>109693186
>I might create a inference/updated training rentry later, this is using the 0.3 Base/Turbo merged model, DiT-only with LM disabled, CFG of 20, only 50 steps (I do not take more steps), MD Storm 4 sampler (which is best), V10 Regrind VAE I linked, everything else is default

I don't know if there's a good dataset for midi + lyrics, ideally already with baked midis. I know there's Hymnery (church song midis, I think baked mp3's? and lyrics - all of old songs, only really old ones), but it's not like ready to go. Even if the results were imperfect, it would be a world first if you got the pipeline laid down (midi+lyrics)-> ai song that follows the melody. I think hymnary is varies, sometimes it's a piano arrangement, sometimes it's the melody? not sure.
>>
>>109693186
>improved VAE

so get
dit/acestep_xl_turbo_Regrind_V1.safetensors
and

wait which vae? lol

I have never played with any of these
>>
>>109693189
Wouldn't it take years generating?
>>
>>109693186
>Voice in real metal songs tends to not be that crisp, listen to some of training data I used
It's not about the voice not being crisp, but I can immediately hear that it's AI made from the voice alone.
The youtube songs you linked have normal voices and sound fine.

>>109693189
Way better voice quality, if it wasn't the pol level usual obsessive shit, I wouldn't have been able to separate that from a real song at first.
>>
>>109693214
the music model is not the same thing as the video model
>>
>>109693189
Because the Minimax dev team have cucked local out of LoRAs and any kind of input audio (covers) for it, likely because it would be very realistic and rival anything available on the cloud, and they also don't want to be held liable for lawsuits. So while it's a sweet improvement to sound quality, their open source promise is questionable. Some are working on reverse engineering the RVQ encoder, but it could take months or years before they succeed. This is best I could do to get music close to Yousei Teikoku from prompt engineering Minimax, and it still does not capture what I want fully (pacing is off)

https://files.catbox.moe/bzuoba.mp3

On top of that, the model is very hard to steer (ACEStep just listens to the prompt, Minimax requires following its prompt templates closely, in addition to not listening most of the time if you stray from that template).

I came across https://huggingface.co/ntc-ai/minimax-music3-concept-sliders
on HF, but nothing supports that yet, and it's still very limited compared to what you can do with ACEStep.
>>
>>109693218
can't find it in the default workflows
is there a good tts too?
>>
>>109693107
can I get a tldr?
>>
>>109693233
all audio models suffer from "out of singer's range = good singer" which also exists in pop music generally. Like Shakura obviously singing a whole song outside of her register.
>>
>>109693234
i don't think anyone made a TTS yet. dramabox is pretty good if you want a TTS model extracted from a video model
>>
>>109693233
I really hate weeb shit.
>>
waah waah spoonfeed me on what kroma should i use for realistic amateur photos
>>
>>109693233
I really love weeb gold.
>>
>>109693213
Not the LoRA, I'm using acestep_1.5_vae_Regrind_V10b-BF16.gguf for the VAE (works fine since I inference on HOTStep.cpp).

HOTStep.cpp is also where you find the experimental MD STORM V4 sampler which I'm using (combines DPM++ 3M and STORK 4 in smart way to get the best of both)
>>
>>109693295
>HOTStep.cpp
sounds cool, I found acestep.cpp has bugs.

does hotstep permit more than 100 steps?
>>
>>109693240
Ctrl +C
>>
If I already feel like I can create whatever I want with Krea 2, what's the appeal of Kroma? Asking in good faith for anons who prefer Kroma.
>>
>>109693339
I did that in the terminal now comfy won't gen anymore
>>
i cant stand non-candid colors, and shit compositions with overcentered subjects of krea, how do people stomach this for realism?
>>
>>109693263
>No gen
>>
[Fun discovery] The fl2va model behaves pretty well at 12 steps with no loras except Mystic V4.
>>
>>109693366
>i cant stand non-candid colors
what the fuck does this even mean? you mean washed out? go get a lora you fucking idiot
>>
>>109693380
that was me begging for help, unironically
>>
File: vase.jpg (55 KB, 1184x656)
55 KB JPG
H3 is so fun

https://files.catbox.moe/vvik1b.mp4
>>
>>109693385
>>109693388
>Unironically nogen
>>
>>109693324
Haven't tested it, but the slider on the UI goes up to 300, and you could possibly theoretically input more. HOTStep.cpp is like acestep.cpp but with many enhanced features. Not sure why you'd want that many steps though.
>>
>>109693389
I miss Paul Rubens :(
>>
>>109693366
I've never had a model respond to composition prompts like rule of thirds, or "no center composition" or center composition in negative prompt. It's fucking weird.
>>
>>109693397
diffusion models are overfit garbage
>>
>>109693397
it works better if you prompt for the subject to be on the side of the frame
>>
>>109693386
Not that anon but is there just one pissy deliberately retarded poster in ldg? I see this on a regular basis here and I don't know of anywhere on this site that's quite like this. You never know what's going to set this aspie off. He doesn't understand what "colors like in a candid photo" means but the E S and L letters are worn off his keyboard. Stop posting.
>>
>>109693417
What the fuck does "non-candid colors" mean? Explain it to me because it sounds absolutely retarded and made up. Googling it returns pretty much nothing so I am going to assume it's MADE UP.

Gemini gives me this:
No, "non-candid colors" is not an established art, design, or photographic term.
>>
spent 2-3 hours intensely studying benedikta's appearance, design, captions and cross checking and researching on captions to make sure there right and on point. I'm pretty sure I've nailed everything down and now its time spend the next 8-15 hours doing the captions for the images. I really want to perfect this as good as the lunafreya and namine but avoid it being janky and rigid as raven and aranea highwing. Rank 64 has better chance of capturing the graphical render style and visuals but the default clothing elements may start to bleed on to other gens even when not prompted for. Going to add high quality synthetic images to dataset of her in different outfits just to play it safe. wish me luck.
>>
>>109693433
>white
>>
I know this is LDG but If there was a god, Wan 3.0 would have been local and open source... the amount of things I've genned so far of my oneitis is insane.
>>
>>109693430
You confidentially asserted what you thought it meant and feigned ignorance just to get mad about it. Nothing in your chain of motivation and action makes any sense, nor would it make sense for me to discuss the thing you decided to sperg out on. My entire point is you should stop sperging.

Also interestingly enough you are the second anon this week to use a contextless gemini query in a reply, and that anon was also a massive retard. I hope this isn't the next gen alpha trend.
>>
File: 1771074641660947.jpg (62 KB, 976x850)
62 KB JPG
>>109693451
the cloudberg will dox you and put you in jail for this crime
>>
>>109693451
Why are you bringing up god in relation to chinkware

>>109693454
Shut up you brown retard. I'm not reading all that brown fingered retard shit. Gemini says you're dumb as fuck, by the way.

The word candid means spontaneous, informal, or unposed (most commonly used in photography, e.g., "candid camera"). Because "candid" describes the nature of a capture or moment, applying it directly to colors creates a mismatch of terms.
>>
>>109693433
>the default clothing elements may start to bleed on to other gens even when not prompted for
doesnt blank prompt preservation or one of the other knobs in ai-toolkit fix this?
>>
>>109693463
>Because "candid" describes the nature of a capture or moment, applying it directly to colors creates a mismatch of terms.

Is that why you typed "you mean washed out?" with your neovagina blood?
>>
>>109693451
all wan iterations after 2.2 were barely an improvement if at all, we would need wan 5.0 in order to compete with h3
>>
>>109693391
It's not always beneficial, but it can help with complicated lyrics.

does hotstep.cpp (I know it's unlikely) have an extend and patch system? Because imo that's how old Udio worked - it automatically sewed your song back together, but extend was just repaint with future like we do with ate step, but also with some past lopped off. but to do this, it can be a little tricky because the vae messes around a little (not much) with the unrepainted audio. I do this in Audacity, but this method is way faster for building songs imo.
>>
I lost my neovagina in neotokyo
>>
File: h3_00011-e.webm (709 KB, 640x1136)
709 KB
709 KB WEBM
any ideas why my h3 vids get bleeding edges at the bottom or right? tried changes many things but no luck.

seems to take color from the opposite edge
workflow and references:
https://files.catbox.moe/uq2b8b.png
https://files.catbox.moe/x3juju.mp4
https://files.catbox.moe/f71ify.png
>>
>>109693473
I hated the 2.2 double model thing, even if I get the intent. The fact I had to get twice the loras was so annoying.
>>
comfyui and adobe are holding initial talks. just to get you really excited
>>
>>109693520
>comfy sells it to adobe
>>
File: 00033-4293344848.png (2.23 MB, 1792x1152)
2.23 MB PNG
>>109693466
not sure what that is, is it caption dropout? i set that to zero. this is what i mean by bleeding details with aranea highwing. The raven(stella blade) lora with original rank 64 and the 2nd attempt at rank 32 suffers from this 85% of the time. Lily lora at rank 64 is way more stable and balanced than raven.
>>
>>109693526
>she has two pcs
>she's just like me!
>>
I feel euphoric. I didn't even wear my fedora today, is SHE thinking about me? is SHE digging me?
>>
>>109693523
>Adobe buys it
>Sas
>Ignores licenses and starts purging forks from the gun3 days
Couldn't happen right?
>>
>>109693593
calm down and vibecode
>>
Some redditor put the FastVideo H3 adapter into a comfy format and I tested it out with the Stacy 40k gals. It actually looks ok, I'd say better than the lightx2v lora in terms of fry, although obviously the dialogue order got messed up. I'm not sure that's the lora since the original gen didn't get it right either.

It's definitely fast at least. I don't know the gen time on the original (which was done with no speedup besides kitchen attention) but this feels a bit faster than 8 steps of lightx2v.

https://litter.catbox.moe/bd49cg11vpzddv5t.mp4
>>
File: debo_mcn_k2_00163_.png (3.48 MB, 1536x1280)
3.48 MB PNG
>>
>>109693510
>any ideas why my h3 vids get bleeding edges at the bottom or right?
the output resolution isnt divisible by 32
>>
File: 1768995604297150.jpg (329 KB, 1600x832)
329 KB JPG
>>109693107
neat lmao, good work anon
>>
>>109693667
What is the best number for that and what does it do?
Did 4chan die? Tried to post for 10 minutes
>>
>>109693670
It's just hallucinating details that aren't there. There's no point to AI image restoration
>>
>>109693676
the resolution selector node in the default h3 workflows has it baked in and most image resize nodes allow you to select the multiple of when resizing, resize everything to a multiple of 32
>>
>>109693488
>I do this in Audacity
Not sure what exactly you mean by extend and patch, but in HOTStep there's a feature to remaster songs with Stable Audio, seperates the instrumental from the vocal to enhance the instrumentals with SA3 by isolating the vocals and extracting it from the track. I have not tried this feature, but its inner workings are interesting, since at the end of the chain it feeds an acapella-enhanced vocals back to the Stable Audio enhanced track. Worth a look if you're interesting in remastering the audio ACEStep gives you, though it's probably not perfect.
>>
>>109691997
Exquisite. Now if only her tiddies were out.
>>
>>109693691
ok, it's like this.

create a 45 second clip

ok

now we extend it. um. I think Udio included the whole thing in "extend". so that means a 90 second latent now, this in ace step is "Repaint" where like you set it to start at maybe 44.8 seconds (might need more overlap, not sure), and end at 90 seconds. It will extend. Repaint is the extend feature, either way, into negative works too (ie -45 start and 0.3 end, results in a 90 second output).

The tricky part is next. Trim the audio to only have the last 45 seconds, and you extend again!

Now, Udio, I think actually had 135 seconds or so, in retrospect, because it was verrrrry prone to forcing a chorus, but my memories aren't too strong, others may correct me.

anyway, at some point you are going to need to glue it back together. I hope this makes sense.

cut the audio to keep only the end, extend this cut audio to add more to the end, you keep doing this but you need to reassemble the pieces.

reassembly in audacity is like you need to remember where you cut, and you line it up, what you do is invert the phase of one of them, then, when perfectly aligned, playback is silent - but honestly alignment is never quite exactly perfect so there's suuuuper quiet audio.

anyway, the seam may wind up being like... a little pop. maybe. You can repaint that or fix it with non-ai. but usually adjusting the join point does the trick, you have a lot of overlap to work with.

I'm genning or I'd immediately double check my theory about negative repaint, I thought I did that. anyway, I'll test it shortly.
>>
>>109693681
>There's no point
Are you sure, have you checked all the possible points to make sure?

>>109693670
this looks funny :)
>>
File: h3_00013-e.webm (783 KB, 640x1120)
783 KB
783 KB WEBM
>>109693685
thanks man. fucking with samplers and models and shifts etc etc and it was something so simple.

now on to real projects.
>>
>>109693759
What prompt be for this?
>>
>>109693759
no problem had the same thing happen to me after manually inputting a resolution and then saw it get fixed after returning to resize nodes, not trying to go full schizo now but its just interesting to ponder how i was here to read your post of all times on something i dealt with, and im pretty sure nobody here would have even replied. crazy stuff.
>>
>>109693778
hello sir
>>
>>109693759
creepy
>>
File: Capital of Conformity.jpg (21 KB, 336x595)
21 KB JPG
>>109692672
I mean ... unless it reaches a point where the model has a complete understanding of your own psychology (or people's in generals) there remains a gap between "what you want from the model" and "what you get from it". Closing that gap is in part what prompting is all about and if "what you want" is something others would appreciate as well and you are skilled at obtaining it then, yes, indirectly people will come to know and appreciate (You) for your prompting skills.

Take Aze Alter's work. Did he "directly" create his videos in a traditional sense? No, he used AI and acted as an intermediary between "what he wants" (his artistic vision) and "what you get". Maybe you personally appreciate his stuff, maybe you don't but the point is that some people do and it is, in effect, for "his prompting skills" (in part anyway).
>>
>>109692605
Nice Portal 3 concept art.
>>
>>109692759
Modern, innit?
>>
>>109692736
Clear Nihei vibes were it not for those fat tiddies.
>>
>>109692850
Man, this is so good on a technical level. Color, "brush"strokes, bokeh, lighting, art style ...
>>
>>109693759
uncanny valley shit
>>
>>109693882
>>109693882
>>109693882
>>
>>109693746
Yes, I always prefer the non-restored image. Adding details that aren't there is uncanny valley.
>>
>>109689804
>>Suno AI's source data leaked
>https://archive.org/download/suno-ai
ummmm yes plz? so i can get past the stupid censorship and improve copyrighted songs?
>>
>>109693063
YWNBAW
>>
>>109693931
It's funny though
>>
>>109692983
>legible signature
wut
>>
>>109693644
>giga goblin
Oh ... oh no ...



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.