[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


My Beautiful Angel Edition

Discussion and Development of Local Image, Video, and Music Models

Previous: >>109530089

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109529779
reposting gravity gun prompt
https://litter.catbox.moe/47prh52lfk66z942.txt
>>
should make bigger collages. saw some gens that deserved more credit
>>
>mfw Resource news

08/11/2026

>LTX-2.5
https://ltx.io/model/ltx-2-5

>LIGHTX2V v1.0 4-step/8-step Turbo Minimax H3 loras
https://huggingface.co/lightx2v/Minimax-h3-Turbo/tree/main

>ComfyUI-DoRA-Dynamic-LoRA-Loader
https://github.com/xmarre/ComfyUI-DoRA-Dynamic-LoRA-Loader

>DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models
https://github.com/chenhsiu48/DeepFreqMark

>Staying True to the Origin: Continuous Image Stylization with Smooth Transitions
https://reychiaro.github.io/StyleController

08/10/2026

>H3 Motion Context: Chain MiniMax H3 clips
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

>MiniMax H3 Turbo — ComfyUI 4-Step T2V and I2V LoRA
https://huggingface.co/joyfox/MiniMax-H3-Turbo

>HRDiT: Training-Free High-Resolution Image Generation with Off-the-Shelf Diffusion Transformer Models
https://github.com/zylwithxy/HRDiT

>RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
https://github.com/LukieLuu/RoRA

>AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
https://huggingface.co/collections/Apryle/avcap

>Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
https://github.com/cau-hai-lab/PORTA.git

>Unsloth Minimax H3 GGUF (Q2:Q8)
https://huggingface.co/unsloth/MiniMax-H3-GGUF

>MiniMax-H3 for Apple Silicon - Rebuilt from the official weights
https://huggingface.co/uetuluk2/minimax-h3-mlx-rebuild

>Soran’t: Small Next.js front end for video generation on ComfyUI
https://github.com/pwillia7/ai_video_fe

08/09/2026

>Kroma v0.2 — Krea 2 fine-tune (full model)
https://huggingface.co/lodestones/Kroma

>krea2-turbo-bbox
https://huggingface.co/jimmycarter/krea2-turbo-bbox

>Kroma v0.2 Quant
https://huggingface.co/silveroxides/Kroma-Quant/tree/main

>Spectrum for Ideogram 4
https://github.com/Nif00/ComfyUI-Spectrum-Ideogram4

>ClipProj-MiniMax
https://huggingface.co/NicoLab28/ClipProj-MiniMax-H3
>>
File: file.png (2.54 MB, 1216x1632)
2.54 MB PNG
yeah w4a8 quant of ideogram4 uncond works really well, no perceivable difference so far. Also had another gen that i liked better lol
>>
>mfw Research news

08/11/2026

>Alpha as an Efficiency Signal: Visibility-Routed RGBA Image-to-Video Generation
https://arxiv.org/abs/2608.09355

>SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision
https://arxiv.org/abs/2608.09097

>In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation
https://arxiv.org/abs/2608.09244

>Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
https://arxiv.org/abs/2608.09529

>Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation
https://arxiv.org/abs/2608.09385

>DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation
https://arxiv.org/abs/2608.09637

>LogiShot: Logically Coherent Cross-Shot Video Generation
https://arxiv.org/abs/2608.08820

>High-Quality Exposure Correction with Diffusion-Based Image Generation Priors
https://arxiv.org/abs/2608.08720

>IDATA: Scalable Invertible Diffusion for Unrestricted Adversarial Transfer Attack
https://arxiv.org/abs/2608.08734

>When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution
https://arxiv.org/abs/2608.09133

>Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation
https://arxiv.org/abs/2608.09594

>Diffusion Image Editing via Asynchronous Token Decoding
https://arxiv.org/abs/2608.09322

>Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective
https://arxiv.org/abs/2608.09057

>UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
https://arxiv.org/abs/2608.08676

>ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models
https://arxiv.org/abs/2608.08060

>RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation
https://arxiv.org/abs/2608.09226
>>
>mfw MORE Research news

>Unveiling the Secret of AdaLN-Zero in Diffusion Transformer
https://arxiv.org/abs/2608.09438

>Wiener Representation Filtering for VLM Hallucination Suppression
https://arxiv.org/abs/2608.08167

>BAG: Budget-Aware Gating for Diffusion Caching
https://arxiv.org/abs/2608.09231

>MemeMind: Reference-Guided Trace Construction for Offline Context Optimization
https://arxiv.org/abs/2608.09316

>VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
https://arxiv.org/abs/2608.08622

>DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models
https://arxiv.org/abs/2608.09233

>Foundation Models are Implicit Deepfake Detectors
https://arxiv.org/abs/2608.09427

>How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems
https://arxiv.org/abs/2608.07861

>Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression
https://arxiv.org/abs/2608.09176

>Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification
https://arxiv.org/abs/2608.09512

>MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling
https://arxiv.org/abs/2608.08553

>LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
https://arxiv.org/abs/2607.16305

>SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision
https://arxiv.org/abs/2604.06247

>Predict to Skip: Linear Multistep Feature Forecasting for Efficient Diffusion Transformers
https://arxiv.org/abs/2602.18093
>>
Anyone else feeling second hand embarrassment for ltx?
>>
>>109531306
Upload the quant?
>>
>>109531316
Their latest ltx 2.5 latent upscale is blazingly fast. Might be useful for Minimax
>>
>>109531316
sink or swim
>>
is there a lora for mayli for krea? I tried looking but I don't see juan
>>
/trash/ needs help setting up minimax h3:

>>>/trash/84651944
>>>/trash/84652113
>>
>>109531316
i don't know. haven't tested it yet. let's see if they made it soulless by loading it with 3d rendered data
>>
>>109531298
You know the thread is healing when people are complaining about the collage again.
>>
>>109531298
Post one of them
>>
File: 17987835498458556.gif (341 KB, 220x184)
341 KB GIF
>>109531379
>>
>>109531322
doubtful since it would mean loading and unloading a whole ass model / vae / te, just to upscale.
>>
>>109531396
damn, this is pretty good son. t2v H3?
>>
>>
did krea get any relevant nsfw loras yet
>>
thanks for bakering
>>
>>109531428
snofs is the best by far
>>
>>109531449
See, I think it's overbaked and suffers from sameface. I like realism engine a lot more.
>>
>>109531428
realism engine
snof
mystic xxx
aio nsfw
just to name a few
>>
may's assets
>>
>>109531451
>overbaked and suffers from sameface
Never seen sameface, but I guess it depends on who/what you're prompting
>>
>>109531321
heres the conversion script with a readme: https://files.catbox.moe/udgu8p.gz
your md5 for w4a8 cond should be 457e4ad22fabe5df5e84f0c354b948f0 after quantizing int8 release from comfy's huggingface
runs at 3.2s/it on a 9070xt at 1.5mp, make sure to use comfy kitchen attention, can go even faster with https://github.com/Nif00/ComfyUI-Spectrum-Ideogram4 in theory but i havent tested it extensively
>>
Repostan from last thread:
Someone posted a start-up guide on another board that I thought I'd follow.
I'm an AMD ride or die so my experience with local gens is next to nil, but I heard there's workarounds these days. I followed the guide through installing ComfyUI, an interface, a model, and finally a LoRA trainer (osiris) which stonewalled me with ROCm. I've generally wrapped my head around it despite a few questions, but the biggest hurdle is that it's saying it requires Windows 11 and I'd rather put my nuts in a vice and send the winch flying than "upgrade" to that garbage.
>inb4 get linux already
Good idea, but the only distro that's said to work with ROCm is Ubuntu and that's not something worth the hassle.
So the question here is, can I make ROCm work with W10 without complaining, have Osiris recognize it, and keep working, or do I have to set up a VM/Dualboot of Ubuntu just to set up ROCm and gen things locally?
>>
>>109531268
may prompt? how do you get her so accurate across multiple shots?
>>
>>109531447
she's cute

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
I pulled.
>>
>>109531449
link?
>>
Holy fuck updating my shit made a huge difference. .5mp at 15s would just brick my shit even with the turbo lora and now I can gen it as fast as I was doing .3mp before. Anon who helped me with the install scripts, you a real one. Kinda feel like a moron for not doing it earlier but at least I can gen now
>>
>>109531459
to each their own.
>>
>[INFO] Native ops: asym_w4a8_int8, int8_tensorwise, convrot_w4a4 , emulated ops: float8_e4m3fn, mxfp8, float8_e5m2, nvfp4
This looks new?
>>
steamed LTX

https://files.catbox.moe/03pjmw.mp4
>>
>>109531494
Same tactics as last year, yeah he never changes
>>
>>109531495
What update ???
>>
File: morepiratepost.mp4 (3.11 MB, 1056x608)
3.11 MB
3.11 MB MP4
>>109531416
t2v h3 prompt here >>109531295
>>
Poor debo, still too poor and dumb to use minmax so he needs to spam for attention again.
>>
>>109531512
comfy updated his kitchen

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109531520
Oh so you did use Comfy you jackass. Imagine what you could do with it instead of shitposting in here
>>
File: kekekekkeeeeeeeeek.png (1.17 MB, 864x1184)
1.17 MB PNG
what is bruh doin nga naaahhhh JSID already its over keeeeeeeeeeeek
>>
>obsessed with will smith eating spaghetti
>black
>hates his name being mentioned (the rentry)
now hear me out...what if...
>>
Why doesn't Krea support editing?
>>
>>109531540
It does actually
>>
/adt/ also adding his rentry (rightfully) really broke him kek
>>
>>109531544
nobody likes him. he's a crappy person.
>>
File: 1763049270513251.png (982 KB, 1056x1056)
982 KB PNG
Created a bunker thread on catalog
Just search "Mimimax" on the catalog.
This fucking shit is unusable
>>
>>109531556
nice try.
>>
File: 1784582810466641.jpg (202 KB, 2048x1536)
202 KB JPG
why the fuck are you genning on windows?
>>
>>109531428
Why a lora instead of a finetune, assuming you mean a general nsfw lora.
>>
>>109531566
Because I'm not gonna dual boot and I only have one pc. I know this is /g/ but I shall not install gentoo
>>
>>109531586
why would you dual boot?
>>
>>109531586
What do you do that can only be done on windows? 90% of games run on linux now. and better than on windows.
>>
>>109531516
thanks, that's pretty good man. did you use llm or are you just this good?
>>
when you swap image references in the reference workflow

https://files.catbox.moe/7gwsbz.mp4
>>
>>109531590
beats me. he has to play fortnite maybe
>>
>>109531573
well I assumed there isn't one yet, if someone has done a full fine tune on krea2 already, sure share it
>>
>>109531586
I've been running linux for nearly two decades. Just switch already, anon. It's so much cozier.
>>
>>109531592
I'm a lolbab sadly
>>
>>109531566
what do I gain by switching to trooonix?
>>
>>109531601
>It's so much cozier.
for real. my life is so much better now that my gaming PC runs linux.
plus you don't even need to learn how to use it anymore. LLMs can solve all your issues, and help you customize your system exactly the way you want it.
>>
>>109531604
>malware enjoyer
oh... my condolences.
>>
>>109531615
freedom
>>
>>109531600
You realize that even if you get the links banned, we'll just make new ones?
>>
I was promised faster gens from pulling and I did not get them... I want a refund.
>>
>>109531592
>90% of games run on linux now. the only ones that dont work are online games that have been intentionally blocked on linux. linux probably supports more windows games than windows itself due to how good it is at playing old outdated games.
>>
>>109531635
you think you can reason with someone spamming, derailing, and trolling every single day for years on end and never leaves his computer except for chicken nuggies?
>>
>>109531635
He's not the brightest lol.
>>
>>109531647
nta but to warn newfrens
>>
>>109531651
if riot ever fixes their shit I'll do it. I just like pushing stupid pieces around in tft but now I just gen I don't even play vidya
>>
>>109531295
Good gen.
>>
people are skeptical of ltx 2.5

https://files.catbox.moe/hahxva.mp4

>>109531720
kill yourself you spamming faggot hope your mom gets cancer
>>
>>109531361
And why should I, a gigabrainchad, help them?
>>
>>109531733
ok simpsons miku got me. this model is so charming
>>
>>109531650
>he didn't backup his installation
>>
>>109531761
>git reset --hard HEAD@{1}
>>
>>109531752
mental illness, spamming a 4chan thread all day because you have nothing of value going on in your life

find a bridge and jump
>>
>>109531733
>>109531780
Just ignore him anon he wants you to reply and notice him because he does not get any attention IRL
>>
>least obvious samefagging
>>
are people thirsty for a namine, aqua and larxene character loras for krea 2? cooking up kairi (KH1 2002) lora at the moment.
>>
>>109531419
BASED
>>
>>109531803
is wanschizo the anon that calls other autists and is generally very rude?
>>
>>109531814
that's goo anon, he's extremely racist.
>>
>>109531782
this troll poll spawned with 5 yes votes in the first minute
>>
>>109531809
I would be down. I just hope the loras are good tho..
>>
File: revealedinadream.png (1.76 MB, 1024x1024)
1.76 MB PNG
the actual truth. which group has lots of free time (since they dont draw)?
>>
>>109531300
>>109531310
thanks!
>>
File: 1781992226163916.png (401 KB, 721x718)
401 KB PNG
image model NOW NOW NOW NOW NOW NOW
>>
>>109531840
The summoners go INSIDE the circle (which protects them). The things summoned are bound to a TRIANGLE OUTSIDE THE CIRCLE.
>>
He might actually make me buy a pass.
>>
May I ask someone for Krea 2 workflow?
>>
>>109531861
you may
>>
>>109531856
imagine this being your only purpose in life
>>
>>109531843
I used to run his leaks through img2img and tards on b begged for more of le hot woman
>>
File: 1781337643790085.webm (1.48 MB, 1104x1488)
1.48 MB
1.48 MB WEBM
What a trash thread
>>
>>109531885
Then improve it.
>>
Alright weebs, shill me on krea 2 and the anime capabilities. After using MiniMax H3, I'm on board with this bigger model = better concept.
>>
Any good loras for minimax yet? Is it still not possible to get good panties?
>>
>>109531885
Got any more tio-tots saved?
>>
>>109531905
>Any good loras for minimax yet?
No
>Is it still not possible to get good panties?
You don't need LoRA's for stuff like that when you can use reference images/videos
>>
>>109531905
ref model made me realize lora concept is shit, and they’ll always be shit fucking up everything else.

learn2ref
>>
>>109531836
it would be not and it will be good as usual
>>
>>109531868
Ok, Might someone provide me one then?
>>
>>109531885
seems like the talisman links do not work lol
>>
h3 reference, best model all years

pay attention /ldg/:

https://files.catbox.moe/a1kjlk.mp4
>>
feels like its just some obsessive need to be the baker of every AI thread
if his hijack gets noticed and challenged he tries splitbaking/early baking in an attempt to get the original baker banned
then if that doesnt work he trolls until enough people leave that his next hijack is successful
>>
>>109531923
No, but thanks for asking
>>
>>109531920
I have transcended
>>
https://rentry.org/ranfaggot
why does this one never get added?
>>
>>109531903
for anime I still think noob/illustrious are the best, like nova anime or wainsfw. fast gens, tons of knowledge, and controlnets.

for editing, klein edit 9b distilled
>>
>>109531861
Unironically use the default. What's wrong with it?
>>
>>109531944
For mixing styles, yeah. Outside of that is Anima territory.
>>
>>109531920
the problem with ref is that you need the ref and it might look like the ref
sometimes you want just similar not the exact copy
>>
>>109531935
why does every single ai voice still have that distinctive robotic sound to it? i would never guess that voices would be this hard to replicate
>>
>>109531903
Anima is SOTA for drawfag anime and Krea is SOTA for illustrations .
>>109531944
Antiquated and outdated model linage.
>>
>>109531958
its a you problem, keep learning
>>
File: file.png (2.01 MB, 928x1104)
2.01 MB PNG
>>109531295
Portal gun but it's Mudeng.
https://litter.catbox.moe/tpzh5j.mp4
>>
Thoughts?

>>>/vp/59499365
>>
>>109531969
it's kinda funny how on point it is tho lmao, it's even in anis rentry.
>>
>>109531977
Why are you linking others gens to throw us of the scent? We know who you are.
>>
What is RTX upscale? Is that something different than the video setting in my nvidia control panel?
>>
>>109531944
Explain why you have illustrious over anima.
Anima can gen higher resolutions.
>>
How to apply Kitchen Attention Nodes ??? I heard it has same speedup as Sage but with no quality loss
>>
>>109531981
???
take your meds
>>
>>109531984
>Explain why you have illustrious over anima
He is unable to tell that mixes and merges have their own strong slop style, he can't tell the difference between the most generated looking image and a true trad kino soul drawing.
>>
File: 00073-4156341952.jpg (370 KB, 2880x1728)
370 KB JPG
>>109531903
beautiful and powerful text 2 image model being held back because vramlet weebs don't want to train character loras for it. Majority of the soo called vram chads with xx90 series cards disappoint me with their laziness.
>>
>>109531960
Bad samples.
>>
>>109531982
It's their own up scaling method instead of another upscale model.
>>109531990
Alice would never be so rude, you just confirmed it.
>>
>>109531996
why are you not making any h3 gens, sasori?
>>
Is there a nsfw /ldg/ somewhere?
>>
>>109532010
/hdg/ on /h/
>>
>>109531966
it's your problem if you think that ref is a replacement for lora
gen more and you might find out why
>>
>>109532010
There's one on gif for only local models.
Most other yellow boards have their own too. Check on e for edg
>>
>>109532010
just stay here, people catbox nsfw shit all the time
>>
>>109531960
does this sound robotic?

https://n.uguu.se/bQdalMdE.mp4
>>
>>109532017
but 90% of minimax loras on civit are cope?
>>
>>109532029
Oh. It's you again. Great.
>>
>>109532003
I never proclaimed to be this "alice".
Take your meds you crazy spastic. Being terminally online is rotting your brain.
>>
>>109532029
I'm gonna BLACKED her too if you don't stop posting the same gens over and over again.
>>
>>109531987
Bro..........................................................Pls respond..................................
>>
>>109532033
Anon, why haven't you taken your medication? It's dangerous for you to remain unmedicated.
>>
>>109531971
basado
>>
File: MiniMax_H3__00024-small.mp4 (3.54 MB, 1120x1120)
3.54 MB
3.54 MB MP4
>>
>>109531996
> vramlet weebs
Not all of weebs are blacks living on gibs in the USA.
>>
>>109531984
illustrious + tag autocomplete + controlnets can do basically anything and the character/anatomy knowledge is very good

also, speed: it's super fast at high res or upscaled
>>
>>109532038
Someone asked why voices on minimax always sound robotic, and I provided an example to the contrary, complete with a query.
I really couldn't care less what autistic little retards like you do. I'll gladly keep posting whatever upsets you.
>>
>>109532038
based jew
>>
>>109532030
yes unfortunately
need more time I guess
>>
File: 1757263783040111.png (57 KB, 212x212)
57 KB PNG
>tell deepseek to generate a smut prompt in a code block
>spam click the copy button as it writes the prompt
>copy the prompt before the deepseek removes and censors it with the "i cant let you do that dave" message
>>
>>109532051
also, you can take an illustrious gen and then edit it with klein edit 9b, or even krea 2.
>>
>>109532040 #
don’t do it, it trippled my gen times on my 5080
only helps with 3xxx and older gpus
>>
>>109532065
>not leaving the "i cant let you do that dave" in the middle of your smut prompt
missing out on edging kino
>>
>>109532070
>or even krea 2.
Explain how.
>>
File: b.mp4 (2.41 MB, 1184x896)
2.41 MB
2.41 MB MP4
>>
>>109532076
krea 2 has an edit lora/workflow, although for anime klein edit works best if you wanna edit stuff (outfits, hair, posing, etc)
>>
>>109531971
very nice. the eventual nsfw finetunes will be neat.
>>
Is there a way to generate pixel art that doesn't have the slop feel?
I keep seeing games with AI generated assets on steam get tens of thousands of wishlists. I want to compete.
>>
>>109532078
fuck off debo
>>
>>109532087
Don't do it for the money. Do it because you enjoy generating pixel art cunny.
>>
why is cat jack generating BBC porn?
>>
>>109531903
krea is actually good for prompt comprehension
anima could have easily been better on that side if they just gave it a bigger text encoder which for some reason they didn't
>>
the reference h3 model is amazing.
>using the voice of <Audio 1>.

https://files.catbox.moe/mykgte.mp4
>>
>>109532082
>krea 2 has an edit lora/workflow
Where? No comfy template for "Krea edit".
>>
>niggas who only generate 1girl, standing talimbout prompt comprehension
lol
>>
>>109532116
https://huggingface.co/conradlocke/krea2-identity-edit
>>
>>109532120
That workflow sucks on ice
>>
>>109532117
you don't understand! it needs to understand concepts like BIG BREASTS and POOPING ANUS!!!1!
>>
>check reddit for ltx 2.5 gens
>cat gen
https://www.reddit.com/r/StableDiffusion/comments/1vlwfvj/ltx_25_on_306016gb_ram_05mp_10_second_video_took/
>talking head gen
https://www.reddit.com/r/StableDiffusion/comments/1vly93b/the_simple_secret_to_better_ltx_25_results/
>not a single gen on /ldg/
total humiliation. it would've been better if they hadn't released this model at all
>>
https://litter.catbox.moe/fcxa7p.webm
>>
>>109532117
i'm not someone with brown iq that only knows how to prompt for the most basic stuff
>>
>>109532143
trump should force them to remove USA from their chart. that shits so embarrassing its basically libel for the US
>>
>>109532143
I'm convinced weather a model is received well or not is just based on the budget they have for jeet spammers. ltx is objectively groundbreaking.
>>
>>109532143
im seeing the same trash being posted in this general
>>
>>109532128
just take the lora and plug it in after your model loader, then prompt edits
>>
>>109532143
realistically, no one can compete with the reference model till there is another reference model with image/sound/video. the base t2v/i2v minimax is already great, but reference makes anything possible without a lora.
>>
>>109532184
what is unique about the h3 reference model? ltx could already use character sheets or audio inputs as references
>>
minimax car cruising

https://files.catbox.moe/32pvfv.mp4
>>
File: MiniMax_H3_00140_.mp4 (1.11 MB, 864x464)
1.11 MB
1.11 MB MP4
RTX 6000 Chads, should I be running with ECC on?
>>
File: 1768098300174314.png (1.21 MB, 900x675)
1.21 MB PNG
be honest. has anyone put themselves and their waifu in a video together?
>>
>>109532178
Nah I tried it and it just didn't work. I love Krea over all but it aint for editing
>>
>>109531996
>>109532100
Any advice for lora training? I realistically only prompt for 3-4 styles on Anima anyway
>>
>>109532204
no, you'd slow it down for no gain, it's not like you're running mission critical stuff 24/7
>>
File: MiniMax_H3__00026-small.mp4 (3.49 MB, 1280x960)
3.49 MB
3.49 MB MP4
>>
minimax needs to somehow ignore the reference image and audio video duration. 15 seconds for video or audio collectively is simply not enough
>>
File: 1772393342424309.webm (3.84 MB, 832x640)
3.84 MB
3.84 MB WEBM
I'm dying from laughter
>>
>>109532228
jesus christ i used to have squish toys like that from the 90s. what the fuck
>>
Who is the bigger cancer, debo or ani?
>>
>>109532235
you
>>
File: MiniMax_H3_00026_s.mp4 (1.01 MB, 928x672)
1.01 MB
1.01 MB MP4
>>
>>109532210
it can do some good edits with realistic gens, but klein 9b edit distilled int8 is the best by far for edits imo

also can uncensor stuff with a lora if you want to.
>>
>>109532235
me
>>
File: 1784293485538756.jpg (44 KB, 848x490)
44 KB JPG
>>109532204
My Allah, it just as expensive as a single family house.....
>>
>>109532275
no one here actually bought one. they rent in the cloud or borrowed one from their job
>>
the voice cloning is so good, especially if you get a clean clip. yes I know it's giorno as image 1, it's just a test. all that you need to clone is "using the voice of <Audio 1>." (or whatever source input)

also, kek it was just a headshot I didnt give a body prompt nor did I say JoJo in the prompt.

https://files.catbox.moe/ceo71d.mp4
>>
>>109531268
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
File: 1765851921433386.jpg (1.09 MB, 4078x3682)
1.09 MB JPG
I tested the simplest H3 optimization nodes to see whether they actually reduce VRAM/RAM usage or improve generation speed:

- MiniMax H3 Chunk FeedForward with the global Sage Attention flag enabled
- MiniMax H3 Mem Eff Sage Attention Patch with the global Sage Attention flag enabled
- MiniMax H3 Low VRAM Attention with the global Sage Attention flag enabled
- No optimization nodes and no Sage Attention flag
- Only the Kijai Sage node, without the global Sage flag
- Only the new Comfy-Kitchen attention path, without the global Sage flag

For every configuration I used the same prompt, seed, duration, resolution, sampler and scheduler + I fully shut down and restarted ComfyUI between runs, and tested each configuration twice.

Test system and workload:

RTX 5090
128 GB DDR4
~1.2 MP output
10-second H3 video @24FPS
20 res_multistep steps with the simple scheduler
H3 t2v INT8 ConvRot model
Same seed for every run
Sampling ram and vram every 250ms

So far, the individual optimization nodes are all very close.
The one result that obviously stands out is any accelerated attention versus no accelerated attention at all. Running without Sage/Kijai/Comfy-Kitchen attention doubled generation time. In other words, on a 5090 (and blackwell in general?), using any of the accelerated attention implementations is obvious.
I did not compare image/video quality, and that would require a much larger set of controlled comparisons anyway.
My current conclusion is that, on Blackwell, the accelerated attention implementations I tested are close enough in performance that quality and maybe stability are more important factors when choosing between them.
The other VRAM-oriented nodes are not really interesting. Low VRAM Attention and Chunk FeedForward don't seem to do much, and stacking all of the optimization nodes together produced the worst outcomes in my tests outside of disabling all accelerated attention.
>>
>>109532293
False
>>
>>109532239
>t. debo
>>
>>109532305
Based tester anon thank you for posting your test results
>>
>>109532298
same prompt but goku

https://files.catbox.moe/8l5jr2.mp4
>>
>>109532314
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
how much knowledge does h3 have of pokemon? can it understand the concept of returning a pokemon to it's pokeball?
>>
>>109532305
no flash attention test?
>>
thanks for jannering
>>109532329
it can do summoning from a poke ball but haven't tried a return yet
>>
>>109532305
post this on /lmg/
this is shitpost thread
>>
File: yeah.jpg (620 KB, 1536x2048)
620 KB JPG
>>109532293
you sure?
>>109532275
can you post that in US dollars, I don't know that yuropoor currency
>>
https://litter.catbox.moe/lnat219wk65n4yem.mp4
>>
>>109532209
Nah just real life girl I know and I put them in sexy outfits with a reference picture.
>>
>>109532352
>you sure?
yes
>>
>>109532336
>t can do summoning from a poke ball
may i see?
>>
>>109532298
>>109532318
how long clips are you using?
>>
>>109532336
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
please maintain thread quality
>>
>>109532316
YW.

>>109532333
>flash attention
No, I got tired of testing at some point, and I had to stop myself from doing all possible permutations or include sol attention for example.
Happy to see other anons testing too though, there are probably optimizations that can be done there I'm missing, and applying to other generations of cards too.
>>
>>109532376
fa is the standard for lossless attention buffs. every other one wipes their ass over the quality
>>
>>109532209
the idea of putting myself into any gen is just too lame
>>
i still wonder who the mysterious shitter could be >>109532138
>>
>>109532372
just got a 5 or 8 second clip from here

https://www.101soundboards.com/boards/1616054-jojos-bizarre-adventure-eyes-of-heaven-playstation-4-part-4-jotaro-kujo-voice
>>
>>109532369
its not great, might need more detailed prompt than just summons whatever
>>
>>109531948
Nothing really. Being more specific, I'm looking for a workflow that deals with the problem of turbo's no variance between seeds. I can use the raw model, but it's kidna slow... 3090
>>
>>109532390
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>
>>109532143
israelis cant into video generation
>>
>>109532394
It's something I guess. Need the pokeball opening and spawning the pokemon accurately rather than having the ball transform into the pokemon.
>>
>>109532352
please get a real cpu fan
>>
>>109532352
If you have that car priced GPU you be sure to take an extreme care for it. cmon dude, no airflow at all
>>
>>109532397
>I'm looking for a workflow that deals with the problem of turbo's no variance between seeds.
One that does not use turbo.

I mean you can get away with it by injecting a bit of noise into the latent - and honestly any lora is going to bring back some randomness as well. Noise injection is probably THE technique for what you want.
>>
File: wrongone.mp4 (730 KB, 736x576)
730 KB
730 KB MP4
>>109532404
can do it opening as well
>>
>>109532305
Now run the same test on a 3060 laptop with 16gb of ram
>>
>>109532415
stop using sol-attn
>>
>>109532352
Looks Indonesian, you can find them around 12500€ nowadays, which is 14-15kUSD.
>>
>>109532401
damn wtf i though those were going to be a bunch of 1girl butts
>>
>>109532415
thanks anon, although the ball seems to have two buttons. Also missing the important part of the pokemon actually appearing haha. im guessing it probably needs a reference unless it's pikachu or something
>>
File: 1767489544062178.mp4 (1.11 MB, 640x672)
1.11 MB
1.11 MB MP4
Not as exciting a gen as I hoped.
>>
>>109532426
that particular clip was meant to be like a pov of the viewer getting captured. there was a joke about it in one of the games I think
>>
>>109532434
>Not as exciting a gen as I hoped.
usually, a prompt issue.
>>
>>109532401
https://www.youtube.com/watch?v=ljXz9r97M3E
>>
>>109532384
I know it's good, but I'm too lazy finding or vibe coding a flash attention 4 compatible node to use with comfy. Sage is already very good from my tests on videos vs dense.
>>
What should I generate?
>>
>>109532440
lol?
>>
Has anyone tried using a video as a reference/weak reference yet? In the times I tried it still ended up just copying the reference into the target video. Still happens even with strong prompt conditioning and ChatGPT revision. GPT thinks it's some issue with the model.
>>
>>109532445
black mario poppin a goomba while hittin a fat doobie
>>
>>109532457
based do this
>>
>>109532434
is it rune scape or run escape?
>>
I love have an image open while I'm typing my video idea into my chatbot. There's something ineffable about this ritual.
>>
File: Tripper.jpg (1.5 MB, 4000x1848)
1.5 MB JPG
>>109532352
get a threadripper. you wont regret it until you realize you dont need a threadripper.
also be sure to use a blower-style fan setup for your gpu and not one that exhausts out the sides straight into a metal wall
pic related
>>
>>109532461
are unescape
>>
>>109532453
Don't believe that message. It was that turtle-brained Yurk Twittlebug that put it there.
>>
File: MiniMax_H3_00030_s.mp4 (2.76 MB, 928x672)
2.76 MB
2.76 MB MP4
Weird glitchy gen
>>
>>109532440
uhhhh what the fuck
>>
>>109532480
Pixel sorting in the prompt? Nice.
>>
>>109532413
>injecting a bit of noise into the latent
will look into this.
Do you need to run raw model with 50+ steps? Almost all workflows for it that I saw have them that high
>>
>>109532405
>>109532408
>>109532469
It's a real case now, that was just a temp setup while I was finishing out my new build
>>
File: 1767849954157117.mp4 (1.82 MB, 736x576)
1.82 MB
1.82 MB MP4
>>
Technically i can make a short 5 min questionably legal porn of this.

But i just dont have motivation to do it.
>>
>>109532493
Btw some anon use the turbo lora with the raw model and prefer it over the full turbo model. But you're always going to be sacrificing elasticity for speed.
>Do you need to run raw model with 50+ steps?
I haven't done much low-step testing, and I also use CFG normalization, so I can't say for certain. I've just stuck with 50 steps.
>>
>>109532519
That sounds like a violation of the user agreement, son
>>
>>109532528
Why's that?
>>
>>109532531
The user agreement says you will gen. You don't get to just not gen after all the resources spent to train the model.
>B-but I'm not motivated enough
Suck it up pal!
>>
>>109532525
ok, thank you. My last question is if you have any workflow that includes inpainting or upscaling? Can't find any
>>
What's the best sampler/scheduler combo for a non-turbo h3 workflow?
>>
File: 1760022226170387.mp4 (1.44 MB, 576x736)
1.44 MB
1.44 MB MP4
>>109532352
>>
File: hunyuan.jpg (155 KB, 800x600)
155 KB JPG
>>
>>109532087
Surely a model designed for 64x64 pixel resolution should be lightning quick and almost entirely free of defects due to the lack of fine detail, right?

>>109532089
this
>>
>>109532519
>of this.
of what?
>>
Bonus H3 cute >>>/wsg/6212531

>>109532562
most i see do res_multistep, euler or er_sde, typically with simple scheduler. not really easy to say which is "the best"
>>
is it possible to use a reference pic in comfyui with illustrious/anima and gen similar looking pictures without a really precise prompt/loras ? a few months ago I tried using ipadapter nodes but it was just bad.
>>
>>109532598
>few months ago
H3 is just over one week anon.
Still older than the taste most here has but still.
>>
>>109532562
for me it's er_sde/beta57
>>
>>109532610
How many steps? I read some cope about 16 and my results were pretty shite, but I'm always interested in anything fast and workable.
>>
File: 1646333888461.png (340 KB, 787x720)
340 KB PNG
I'm compiling a library of anime character reference images.
ChatGPT says H3 R2V will capture a character's likeness best with a full body front view, back view and three-quarter view, so I'm generating them using anime (or wai illustrious if anima doesn't support the original style).

With this, I will not need to create character loras for my minimax gens.
>>
>>109532613
I've had very usable gens at 10steps. but 20 is much better for overall coherence and prompt adherence
>>
>>109532598
IDK. probably with complex controlnets etc.

but really most use the newer edit models and/or just newer models that can take longer more accurate descriptions like krea/ideogram (because then you can pretty closely replicate it even via prompt that a VL or such got for you)
>>
>>109532618
Did the blog factory explode
>>
>>109532143
Cool it with the antisemitism
>>
>>109532143
its not easy to follow up h3, especially since we're now spoiled with ref2va
>>
>>109532143
at least debo has a video model that even he can use
>>
File: 1772514138651193.jpg (16 KB, 205x205)
16 KB JPG
minimax reference has been blessed with the sweet sounds of gilbert gottfried

https://files.catbox.moe/pfiry8.mp4
>>
im testing ltx 2.5 and it completely blows h3 out of the water. the jews won
>>
>>109532640
yeah dude same but I'm not gonna show anyone
>>
>>109532618
Would character sheets work well for r2v
>>
>>109532644
based
>>
>>109532640
yeah totally, dude ah ah, LTX2.5 is soooo good!
>>
File: ComfyUI_08619_.png (910 KB, 1024x1024)
910 KB PNG
>>
forsen except he's gilbert:

https://files.catbox.moe/bozzit.mp4
>>
>>109532197
nta, this is far far more powerful mainly because the model is just so much better and it understands things. Just so much creativity on offer, it even has a feature to remove the backgrounds of ref images and they are then used in the environment you choose. So in a way you might have to put more thought into the ref images for what you want or you could actually leave it more random to create maybe a better gen. I would not like to think how complicated it could get especially with the video and audio ref's as well
>>
>>109532648
According to ChatGPT:

>Short answer: collage vs multiple separate images
>Best option in principle
>If the workflow supports it well, multiple separate reference images are usually better than one big compilation image.
>Why:
>each image stays larger and clearer
>the model can read facial features, hairstyle, outfit details, and proportions more cleanly
>there is less risk that it interprets the collage layout itself as meaningful visual structure
>there is less loss of detail from shrinking many images into one canvas

I tried both, and GPT was proven correct, I definitely got better likeness using individual images than one big compilation image. This was with a real person though, maybe anime is more forgiving.
>>
Where are the loras ???????????????????
>>
>>109532685
she's busy raiding tombs
>>
What about character loras can we still make our own???????
>>
Anons using the minimax skill to enhance prompts, is gemma 12B enough or is 31B the absolute minimum?
>>
>>109532689
ok someone gen anon getting a blowjob from 1996 lara croft
>>
>>109532691
I've been using Fable Ultracode
>>
File: MiniMax_H3__00031-small.mp4 (3.6 MB, 1376x768)
3.6 MB
3.6 MB MP4
>>
>>109532710
Did you gen that on the 9070xt?
>>
>>109532716
nope the 5090 like a real man
>>
>>109532708
And it's ok with nsfw?
>>
>>109532691
uncensored qwen3.6 35b seems better to me.
>>
>>109532691
I use 31B and sometimes it's a bit dumb. You can try 12B but I suspect it will be very shit.
>>
why did my comfyui extension button dissapear? it was floating up there and its just randomly gone
>>
How do I stop the little blip of audio at the start of my gens?
>>
>>109532728
>>109532729
Thanks anons, issue is to juggle that with main generation.
>>
>>109532734
sounds like a prompt issue.
>>
How is the RTX upscale node with real life footage? I've seen people post anime and shit with it, but never real people. worries me a little but uh what's the verdict
>>
>>109532734
ffmpeg -i input.mp4 \
-vf "trim=start_frame=20,setpts=PTS-STARTPTS" \
-af "atrim=start=0.833333,asetpts=PTS-STARTPTS" \
output.mp4
>>
>>109532738
>juggle that with main generation.
Don't tell my local gemma, but I've been using her on openrouter because of this reason.
>>
>>109532741
1.25
1.5 is hit and miss
>>
>>109532747
Yeah, but in that case might as well go with bigger models like kimi k3.
>>
>>109532729
to me either of these is a bit basic to dumb for nsfw, unless you found some tune that added nsfw knowledge

>>109532738
unload/reload either model as-needed. probably not the only model you operrate that way. obviously you can get a bunch of seed variant gens off one prompt
>>
>>109532749
I like that she's extremely cost effective. think I've only spent 2cents over 2 days.
>>
ref model: picture reference, and a photo of the skyrim intro

Forsen is represented by <Picture 0>

the setting is the forest in <Picture 1>. Include the cart in <Picture 1>. Forsen is sitting in the cart beside the blonde nordic man in <Picture 1>.

Forsen says "so, where are we headed mister Nord?". The blonde nordic man says "ah, you're finally awake" in a swedish accent. Forsen says "Oh, i'm just going back to Sweden", as a dragon flies over the cart and blasts it with fire, setting Forsen and the Nordic man on fire.

https://files.catbox.moe/wv51rx.mp4
>>
>>109532752
Which LLM would you recommend?
>>
>>109532759
>Picture 0
I believe it starts at 1 even though the connection says 0, going by the guide
>>
>>109532739
Yeah, I'm asking how I keep fucking it up. I'm putting N/A for both the sounds. Should I just say it's completely silent other than the dialogue from whatever speaker?
>>
>>109532772
can't really help without seeing the whole thing.
>>
>Thread Challenge:
No early bake.
>Hard mode:
page 5
>Impossible mode:
page 8
>>
>>109532729
>I use 31B and sometimes it's a bit dumb
really? I planned on using that one, damn, there is anything better locally and also able to be fed images for reference
>>
>>109532771
yea I fixed it now
>>
>>109532772
I've gotten this issue pretty often and haven't found a fix, so at least I can say it isn't just you. You could try being aggressive with timestamping things at 00:00.00 like other coherent sounds
>>
>>109532777
I mean anons done it before, a long time ago. And a handful of times its fallen off the catalog before a new one was baked.
>>
>>109532766
i like this https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF/

could still be more nsfw trained but it's better than any <32B or so gemma at what i tested
>>
>>109532781
*isn't
>>
>>109532305
Low VRAM Attention and Chunk FeedForward simply allow vramlets to gen longer and higher res videos. how many chunks and "head chunks" you need? I'm not really too sure.
>>
>>109532785
usually in euro times since this is an america general
>>
>>109532781
Big emphasis on *sometimes*
it never fucks up the prompt template. just you can't expect it to gen a masterpiece by just prompting it
"make her show bobs"
>>
>>109532777
Lucky digits have blessed this challenge.
>>
File: 3yer.png (300 KB, 795x525)
300 KB PNG
>>109532798
>>
File: 1785134955893786.jpg (53 KB, 681x1076)
53 KB JPG
>>109532798
Yeah but my results ran counter to my expectations: basically I thought more chunk = less vram peaks, except it's the opposite that happened.
>>
i have lost all my motivation to generate on minmax
>>
https://files.catbox.moe/3v6fjx.mp4
>>
>>109532790
Qwen is way too slopped for my taste.
>>
>>109532618
that combined with instant character voice cloning and longer native gens would be a dream come true
>>
is anyone having trouble getting the ref model to NOT create a bright evenly lit scene? can you do dark shadows and harsh lighting?
>>
>>109532813
retvrn to sovl
>>
i cant get gemma 4 q8 to even follow proper prompt generation syntax what a stupid bitch
>>
For any orange connoisseur fans out there, minimax knows the style of Star Wars: The Clone Wars. Just tag it like that with 3D-Animated, and it copies the style from the show exactly.
>>
>>109532825
youre supposed to use qwen3.6
>>
>>109532825
just buy some prompts on promptbase
>>
>>109532813
You're allowed to take breaks anon!
>>
>>109532828
>>109532832
im just gunna use the blowjob lora without a prompt and hope for the best
>>
>>109532827
how the fuck has no one thrown the "and she was a good friend" meme into H3 yet?
>>
>>109532838
Requires multiple passes for the whole speech, then editing them together later. Also, it was never funny.
>>
>>109532838
I actually planned to, but got busy last night. Specifically a recreation of Disclaimer's comic.
>>
>>109532802
got it, I'll try it anyway, I can't run much locally
>>
>>109532710
delightful. unfortunately for you the gpu only pulls 350w, sometimes a spike to 400w. its nowhere near housefire zone like any 12v hpwr connector
>>
Post kino.
>>
File: MiniMax_H3_00935_2.webm (2.13 MB, 1184x896)
2.13 MB
2.13 MB WEBM
https://i.4cdn.org/wsg/1786502062635557.webm
>>
>>109532822
Just kinda telling it to make things dark and grim and nasty seems to work. I've got one (thanks to my clanker) that generated pretty dark and gritty, it starts:
> Cinematic live-action horror with wet biomechanical textures, dim chiaroscuro lighting, deep shadows, and a cold green-purple color palette.
adjust as needed
>>
>>109532855
making a prompt for the new ltx
>>
>>109532864
very nice.
>>
>>109532864
this could have legit been some cartoon version of the story in the 2000s
>>
>>109532864
You know you can get it to do snape's voice by using <d>[English, in the distinct voice of Alan Rickman as Severus Snape] *text*</d>, right?
>>
File: 00007-1612742747.png (2.98 MB, 1536x1920)
2.98 MB PNG
>>109532864
>>109532881
I would've watched this shit as a kid back then. Man, so we really can just take live action movie scenes and give them a 2d spin huh.. there's some serious potential there.

>>109532886
or just feed it the exact audio from the movie for reference.
>>
>>109532886
Wait, THAT's the right place to put descriptions of character voices? Am I about to experience a breakthrough? Don't you lie to me, anon.
>>
>>109532894
If only there was some sort of prompting guide that explained everything. Imagine if the creators of the model did it, even! That would be craaazzzzzyyy.
>>
>>109530531
>warn newfags
translation: they are there to educate newfags about thread drama concerning anons you are unable to share an ai image general thread with. i.e. shit no one gives a flying fuck about
>>
>>109532893
I'm assuming he was genning via t2v on the fl model not ref model. But yeah, obviously that works, too.

>>109532894
It's how I've gotten it to do voices. I only add the series or actor name in there if it struggles, but usually just [Language, in the distinct voice of Character] works for standalone non ref audio generation.
>>
>>109532904
man I read the prompting guide I guess I just missed that part, I only saw that you used it to write [English] and assumed language was its only function. guess I'm a tremendous faggot
>>
>>109532907
Why would you read it yourself instead of having a robot read it for you?
>>
>>109532893
>I would've watched this shit as a kid back then. Man, so we really can just take live action movie scenes and give them a 2d spin huh.. there's some serious potential there.
with enough autism it's legit possible to make a giant ref2v of a story
>>
is there a ref2va turbo lora yet
>>
>>109532874
Thanks anon. "dim chiaroscuro lighting" definitely seemed to help
>>
>>109532894
literally read the prompting guide for protips
>>
>>109532904
to be fair, the prompting guide never says you can do the bracket english thing
>>
>>109532910
how the fuck do you think I'm implementing it? I blame my shitty clanker.
>>
>>109532913
they all appear to work on it regardless
>>
>>109532875
proof?
>>
File: 00010-2382572119.png (3.2 MB, 1536x1920)
3.2 MB PNG
>>109532911
We truly are in a renaissance..
>>
>>109532930
lightx2v T2V turbo WAN lora worked with I2V wan too, but the quality wasn't near as good and it killed motion. i prefer if a turbo lora was specifically trained on ref2v
>>
>use "MM-H3 - Blowjob.safetensors"
>no prompt
>it doesnt make her give the viewer a blowjob
>it just makes her rub her face in a cute way and then look deeper into the eyes of the viewer
fuck
>>
>>109532938
hot
>>
>>109532719
Yes.
>>
Guys, you need to remember to use the quantized heretic text encoder and the quantized vae. Otherwise your gens will be slow.
>>
Is the new lightx2v good?
>>
Knock knock
https://files.catbox.moe/bi6dcc.mp4
>>
>>109532931
soon
>>
File: 000260.mp4 (1.17 MB, 960x544)
1.17 MB
1.17 MB MP4
>>
I'm animating 2-3 pages of a doujin and it's going well.
>>
>>109532969
Post an example of what you mean.
>>
>>109532975
when it's done.
>>
>>109531814
>wanschizo
Yes
He also likes using smiley faces at the end of his messages.
>>
>"overcast night"
>puts bright shining moon in the sky
>sigh
>>
>
>>
>>109531814
>>109532979
This is him by the way, his posts are extremely formulaic and his writing voice never changes.
>>109532053
>>
>>109532957
heh
>>
>>109532982
did you prompt it with no moon? did you put a comma between overcast and night?
>>
Did ani just return or something?
>>
>>109532999
the luddite schizo on /a/ started getting rekt in the /a/ thread so he came here in an effort to get revenge. he calls ppl "wanschizo" cause he saw ani calling someone that once over on /a/.
>>
You're baiting me.
>>
>>109532957
>gives her a bunch of bubblegum pieces
who the fuck does that? well him i guess.
>>
File: 1767610345035416.jpg (60 KB, 550x366)
60 KB JPG
TIME PLEASE STOP FOR FUCK SAKE HOW THE FUCK ITS ALREADY MORNING AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA

ZA WARUDOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOOO
>>
[Identity Mapping]
Forsen is represented by <Picture 1>

the setting is the forest in <Picture 2>. Include the cart in <Picture 2>. Forsen is sitting in the cart to the right of the blonde nordic man in <Picture 2>. Forsen says "so, where are we headed mister Nord?". The blonde nordic man says "oh shit!" and the cart rolls off a mountain cliff, down into a large pit, and explodes.

testing random stuff

https://files.catbox.moe/ym3n3p.mp4
>>
>>109532989
yeah ok I see who he is now. thanks. He was baking the threads yesterday and would insist on never making a collage then he proceeded to call everyone who complained morons and autists
>>
File: 1786380243637.png (287 KB, 2752x1379)
287 KB PNG
Anon showed how to continue a gen with this.
I'm sure this works fine but how am I supposed to also feed it the audio?
>>
Well this is rare. I got face drift all the time in my gens
>>
This model is great for anyone with one specific fetish that nobody ever does just right.
>>
tourist here. So I hear that LTX 2.5 is fast enough to run at real-time speeds now, is that right? Does that mean I can have an animated digital waifu that can do anything I tell her to now?
>>
>>
>>109533070
lol nah
they did all their testing on insane OP rigs, no one is getting realtime, not in any decent enough to use quality anyway
>>
File: 1785038480208599.jpg (123 KB, 1186x690)
123 KB JPG
Does face reference like this works on Minimax ??
>>
>>109533084
yeah probably, try it.
>>
File: 200100_00001.mp4 (2.75 MB, 704x480)
2.75 MB
2.75 MB MP4
https://files.catbox.moe/9fsr96.webm
>>
>>109532957
>>109533021
skittles giants
>>
>>109533087
Nice.
>>
>>109533087
wtf? I want to play this.
>>
>>109533080
:(
maybe next year
>>
>>109532969
Feel free to share some advice for us(me) who didn't get good results making
>>
>>109533092
coincidentally ive been working on this as a game too i just wasnt sure what to pursue. i know what i want, my cock will guide me there, but whether itd make a good game is a totally different story
>>
>>109533084
it should
>>
>>109533070
>I can have an animated digital waifu that can do anything I tell her to now?
This is not dependent on real-time generation and is in fact possible with current generation models and hardware.
>>
File: debo_dm_k2_00059_.png (2.69 MB, 1872x1007)
2.69 MB PNG
>>109533097
based project pursuer. I hope you achieve your goals
>>
>>109533101
Are you talking about animation engines for 3D models?
>>
>>109533109
i hope i do too. ive restarted my project 8 times be cause i cant stop hitting a wall and lose motivation.
>>
>>109533111
LTX is not that.
>>
>>109533097
just know that nobody will buy it once they learned you genned all the relevant assets and maybe vibecoded it as well
>>
What are you prompting, anon?
>>
>>109533136
kino
>>
File: debo_dm_k2_00060_.png (2.58 MB, 1872x1007)
2.58 MB PNG
>>109533116
I know the feel. easy to achieve anything when you're in the zone then impossible to even open the project when you're not

you can do it, anon
>>
>>109533136
kino as well (but krea 2 isn't cooperating with the lighting)
>>
>>109533096
you should look for the posts from the BLAME! anon hes shared some good advice
>>
>hard mode:
Completed!
>>
File: 1773350931054457.mp4 (423 KB, 768x544)
423 KB
423 KB MP4
>>
>>109533135
i have connections to make it happen i just want to make something myself by hand rather than with slop. i dont want to shop a single slop related thing in my life

>>109533146
i made a bad choice of unreal engine visual scripting (as a c++ project) despite no programming background because its my best shot
>>
>>109533079
What does she feel like?
>>
>>109533150
>hes shared some good advice
Did I?
ps: >>109532969(Me)
>>
>>109533150
>BLAME! anon
Could I get spoonfed? Am sadly not as much online here as I'd like to be so don't want to crawl around the archives.
>>
>>109533100
>>109533086
I still get a face drift with different face expression lol
>>
There aren't many images in existence worthy of being animated it seems.
>>
>>109533096
mostly doing it to learn video extending.
maybe you can tell me what you're having trouble with?
>>
>>109533158
Bags of sand
>>
File: Misty_Dance_00079_.webm (1.92 MB, 832x640)
1.92 MB
1.92 MB WEBM
https://d.uguu.se/ijgLTfnu.webm
>>
File: 1770915735723973.mp4 (827 KB, 768x544)
827 KB
827 KB MP4
>>
>>109533173
>now misty gets a Toxic turn
very kino, nice ass rotation too looks very smooth.
>>
>>109533174
Oh so that's what was over there!
>>
File: h3_00004 (1).mp4 (3 MB, 832x640)
3 MB
3 MB MP4
>>109533164
>>>/wsg/6211923
and vidrel was my first early testing.
>>
>>109533179
Yeah, thought it was time to give May a rest and test a new set of picture references. I think it came out pretty well.
>>
>>109533174
sovl
>>
this next one will be supreme kino just wait
>>
ltx is very nice. i don't need to enable swap space to use it
>>
>>109533173
Wanschizo, does it ever occur to you that you could just post these gens here forever and not bother shitting up /a/'s catalog
>>
>>109533198
bro take your meds already.
>>
>>109533174
https://www.youtube.com/shorts/ffySP5SMfcw
>>
20 hours wasted. Just like that
>>
The difference between "fully_preserved" and "partially_preserved" can be tricky.
>>
File: stich.mp4 (556 KB, 608x1056)
556 KB
556 KB MP4
hmmm... so the video in place helps. but there's a saturation change.
>>
>prompt:
>POV
>with his right hand he reaches out
>...


>video shows him using his left hand
>>
File: 464264.webm (3.09 MB, 420x291)
3.09 MB
3.09 MB WEBM
testing the new multi-shot support in ltx 2.5
>>
>>109533173
pokemon...pokemon...
>>
>>109533244
For a video, it always seems to be fully preserved.
>>
>>109533252
Is that different than the H3 reference model?
>>
>>109533251
"with his other right hand he grabs his left penis"
>>
>>109533254
Yes, debo. Get used to it.
>>
File: 4642642.webm (2.78 MB, 420x291)
2.78 MB
2.78 MB WEBM
>>109533257
ltx is a different model from h3
>>
>>109533258
grabbing penis is never right
>>
>>109533252
idk, this has some kind of LTX sovl.
>>
>>109533258
>honey, why does only one of your cocks get hard anymore, instead of both? do you still love me?
>>
>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109533168
I'm trying to use a black and white Manga so the colours are messed up between frames and since it's low res I can't really give it a good upscale either
>>
File: 2026-08-12_0.webm (3.42 MB, 704x1056)
3.42 MB
3.42 MB WEBM
>>
>>109533275
thanks for the reminder to include it in the next OP
>>
File: debo_dm_k2_00062_.png (2.77 MB, 1872x1007)
2.77 MB PNG
>>109533275
gm
>>
>>109533280
>the colours are messed up between frames
ah see I'm using BLAME! and keeping it monochrome specifically because I didn't want to fuck with that.

Probably want to use your first gen fed as a grounding reference for the colors on future gens.
>>
>>109533265
Idk what gen is what model but this one mogs the other no cap
>>
>>109533271
yeah it makes everything looks real by default since the dataset is less synthetic
>>
>>109533293
It changes scenes and I did try to use image 2 as a color reference but it didn't always like that. Swapping the hair colors and clothing colors between "shots" just ruins it for me.
I would've never imagined a local model being this good and fast when novelai leaked and that was the big thing. Hope the world makes it a few more years for more of these ones.
>>
File: 00010-935572018.jpg (518 KB, 1920x2880)
518 KB JPG
>>
File: 578639842309.mp4 (1.19 MB, 1056x608)
1.19 MB
1.19 MB MP4
>>
>>109533314
damn this looks really nice, like the actual Kh1 fmv kairi. Is this krea 2?
>>
>>109533319
there's so many!
>>
File: 1785604627667703.mp4 (264 KB, 640x640)
264 KB
264 KB MP4
>>
>109533319
missing the flames coming up at the end.
>>
>>109533327
NooooOO be nice!!
>>
File: 00011-2501159376.jpg (342 KB, 1920x2880)
342 KB JPG
>>109533320
yes, literally 80% of images i used to train this kairi krea2 lora were from the cgi fmvs from kh1, kh2 and dream drop distance. the last 20% were just high quality character model screenshots from the hk1. I socked and amazed how it turned out.
>>
>>109533351
How'd you set up your krea training env?
I tried on a clean venv and it just refused to work
>>
>>109533351
glad to see you back sasori
>>
>>109533183
Hope you keep at it
>>
>>109533326
artistic liberty. represents the schizos multiple personalities.
>>
File: 1763974885010605.mp4 (320 KB, 640x640)
320 KB
320 KB MP4
>>109533327
>>
>>109533374
deserved for having ever posted in /s*g/
>>
>>109533351
awesome, god damn. would you share your lora with the class? I have so many ideas i'd love to pull off with this.
>>
Finally manage to preserve the face.... after few hours. Prompt enhancer works and those face reference works. Just a reminder you need an every angle of the faces.......
>>
Has anyone, like, tested the sigma shift, or is it purely vibes?
>>
>>109533395
Everything about this hobby is vibes.
>>
What is the meta for image editing? Say I have an image of a character doing something, and then I have a sample image of another character, and I want to replace character 1 with character 2 but leave everything else intact. What is the best model/tool for that? Would it still be qwen image edit in comfyui?
>>
>>109533399
shut the fuck up
>>
>>109533395
i did testing with it. lower values cause the prompt motions to be followed more aggressively which can cause the video to speed up if your duration is too short for the prompt you used. it's good if you want fast camera motion
>>
is there any way to get speech to speech? minimax seems to just be transcribing dialogue into raw text and having the voice reference read it instead of truly replicating it.

https://files.catbox.moe/fpevv9.mp4
>>
File: 1766297006616774.mp4 (573 KB, 640x640)
573 KB
573 KB MP4
>>
>>109533354
i used ostris ai toolkit, qint8 for both the transformer and text encoder. 5000 steps and 1024 res. Most of the images i used are way above 1024res and are downscaled to the resolution buckets pertaining to 1024 res. High resolution images and precise captioning seems to be key for getting near 1:1 results. kairi is simple character so rank 32 is alright for her but someone like aranea highwing is way too complex and requires rank 64 with many more images of her just to play just get near accurate 1:1 results.
>>
>>109533411
Why? I keep getting told to go to a different general with my question.
>>
File: 634743.webm (1.6 MB, 420x291)
1.6 MB
1.6 MB WEBM
>>109533295
everything is the same model, i'm still figuring out how to prompt this correctly. the multi-shot stuff is not always working for me
>>
>>109533412
By aggressively do you mean better adherence, or just more dramatic and emphasized?
I.e. it'll miss an action in either setting, but when it does the action on a lower sigma it'll be more clear?
>>
>>109533418
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), guiding delivery and speaking rate without copying the original signal.
>>
The longest video i done is 20 seconds. Good job turbo lora.
>>
does specifying over 30 seconds break everything for anyone else?
>>
File: 444352.webm (3 MB, 576x320)
3 MB
3 MB WEBM
>>109533440
the way i imagine it is that each sentence you write in your prompt is a single action. the sigma determines the blending between all of the sentences. if you have tons of sentences with small actions, then a large sigma might average all of it together to be smooth and probably not follow the prompt exactly. lower sigmas will literally execute each sentence you wrote and it can make the video become erratic as it tries to fit everything into your duration. try using the same prompt and seed and generate both extreme ends of the sigma value to see what i mean
>>
The quality difference between 20 and 40 steps is quite noticeable.
>>
File: 00015-147914423.jpg (323 KB, 2880x1920)
323 KB JPG
>>109533382
have fun with it. https://gofile.io/d/shxHnB
these were the tags i mainly used, not necessary to use all of them.
kairi, Kairi (Kingdom Hearts), Kingdom Hearts 1, 1girl, solo, female, girl, youthful, young teenage girl, 14 years old, fair skin, light skin, blue eyes, short hair, red hair, hair bangs, layered hair, hair between eyes, red eyebrows, black choker, necklace, teardrop pendant, collarbone, sleeveless, white cropped tank top, white cropped tank top purple trim and purple straps, black undershirt, yellow sweatband on left forearm, purple armband on left upper arm, yellow and black bracelets on right arm wrist, navel, midriff, dark purple belt, silver belt buckle, light purple short skirt, light purple mimi-skirt, scalloped hem on skirt, flower petal cutout pattern on the scalloped hem of skirt, side slit on skirt, side slit on both sides on skirt, purple bike shorts, purple bike shorts worn underneath her purple skirt, oversized shoes, chunky shoes, white shoes, yellow laces on shoes, blue and purple soles on shoes, slender build, slender legs,
3d, 3d render, cinematic render, high fidelity, cgi, cutscene, videogame cinematic cutscene, cinematic scene, cinematic cgi graphics,
>>
>>109533469
>>109533469
>>109533469

move along
>>
>>109533470
>one of the links removed
i see you, fucker
>>
>>109533474
>>109533474
>>
>>109533463
This is a fantastic description. I have some prompts with a million little things happening so I'm interested to test it a bit.
>>
>>109533470
>>109533477
oh no
>>
Not one inch
>>
>>109533476
just post it on the other thread, I wanna post my gens
>>
so close to page 8 :d
>>
>>109533477
real bread.
>>
new bread

>>109533469
>>109533469
>>109533469
>>
>>109532500
lol
>>
>>109532564
not the anime girl AND the computer
>>
>>109533474
>>109533474
>>109533474
Avoid d*bo thread
>>
>>109533952
>spamming an abandoned thread
mind broken



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.