[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 1786228555489579.webm (3.82 MB, 960x728)
3.82 MB
3.82 MB WEBM
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109513473

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Qwen
https://huggingface.co/collections/Qwen/qwen-image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Chroma
https://huggingface.co/lodestones/Chroma1-Base
https://rentry.org/mvu52t46

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.com/debo
https://rentry.com/animanon
>>
>>109515244
I'd love to see an entire episode of that HP Anime
>>
You know what, I was too harsh on the 0.1 speedup lora from lightx2.
>>
File: t2v_MiniMax_H3_00015_.mp4 (1.77 MB, 672x1216)
1.77 MB
1.77 MB MP4
this shit rocks, what a time to be alive
>>
i keep having to reload the turbo lora. is that normal? can it be baked into the main model some how?
>>
File: three-toed-sloths_thumb.jpg (700 KB, 2048x2048)
700 KB JPG
LTX 2.3 1080p 15 seconds gen time: 40 min
MM H3 1080p 15 seconds gen time: 3 hours
Next version of LTX can mog H3 if it comes with similarly big text encoder and fixes its hallucinations
>>
Anyone experiment with using the i2v as a simple reference model by giving it an image with all your characters then prompting for a cut immediately with the scene you actually want? It works with single subjects, might also work with multiple. Might be good for quick and dirty gens without multiple characters without have to use the slow ass reference model.
>>
File: d.png (46 KB, 1419x749)
46 KB PNG
>decide to click
whats even the point throwing a fit every bake over dead links?
>>
>>109515217
Not bait. The conversation was about using AI to create better software support for non-nvidia GPUs.
>>
File: 14_l.jpg (83 KB, 590x320)
83 KB JPG
>>109515339
soon, goyim. soon...
>>
>>109515342
OP is a troll and replaced .org with .com
>>
>>109515342
its https://rentry.org. the anon changed it this OP
sneaky cunt
>>
File: 756428.webm (1.05 MB, 576x320)
1.05 MB
1.05 MB WEBM
progress so far
integrated_multimodal_description: [Shot 1] 1944. real life. live-action. raw footage. combat footage. go-pro. soft video quality. desaturated colors. highly dynamic and shaky motion. motion blur. the perspective is on the left side of the cockpit of an american world war 2 high altitude heavy bomber (b-17, buttons and switches on the top corner of the frame, 2 narrow front windows, nose of plane hidden by the flight instruments). blue sky. cloud layer below. other similar bombers (brown, 4 engines, single tail rudder, open sides behind the wings, top turret) are flying in our formation. the bombers have contrails. flying through heavy flak fire. our hands and legs (pilot uniform) are at the bottom of the frame holding the flight wheel. lots of analog flight instruments. active combat.
at 00:00:00, the perspective is looking out of the window. we have to lean in order to see the the large brown wing (2 engines) outside the side window. big bursts of black smoke exploding in the sky. one of the flak shells explodes close to the wing and slightly damages it, causing the cockpit to jerk and shake the perspective. the perspective turns to our hands on the flight wheel at the bottom of frame, and then quickly turns to look at our co-pilot (flak jacket, leather cap with goggles) who has a nervous expression hidden by his oxygen mask. the perspective turns to our hands on the flight wheel at the bottom of frame.
overall_soundscape: flak booming with a deep engine rumble. airframe is rattling.
non_diegetic_music: N/A
>>
>>109515321
>the streetlights reflected off the window stay when the indoor lights turn on
kino
>>
File: 163689926459943.mp4 (3.69 MB, 640x832)
3.69 MB
3.69 MB MP4
>>
>>109515387
did she shit her pants?
>>
>>109515392
seems like it
>>
File: willem-dafoe.gif (3.72 MB, 480x640)
3.72 MB GIF
>>109515392
>>
>>109515387
Jesus Dennis

>>109515358
Kino
>>
File: 7564282.webm (943 KB, 576x320)
943 KB
943 KB WEBM
changed it to crack the glass. way more kino. now it's time to push the duration and shoot the plane down
>>
24gb isn't enough

https://files.catbox.moe/sejwv2.mp4
>>
D-did the bots finally die? Are we free?
>>
>>109515358
>>109515434
>How is the audio?
>>
Is there a computer spec vs performance chart for Minimax yet?
>>
>Maintain Thread Quality
https://rentry.co/debo
https://rentry.co/animanon

Always be vigilant for dogshit bakers.
>>
File: 1780695896849001.png (2.44 MB, 1402x1122)
2.44 MB PNG
Since /ldg/ has been saturated with video generated slop, we need a new general for images.
>>
>>109515587
>we need a new general for images.
What would that accomplish?
>>
File: 444541238234605.mp4 (3.87 MB, 544x960)
3.87 MB
3.87 MB MP4
>>
>>109515572
it mostly sounds like movie sound effects. h3 hasn't been impressive in the sound department
>>
File: file.png (5 KB, 296x51)
5 KB PNG
>>109515583
I see no namefags. Everyone is equally retarded to me unless proven otherwise by their posts.
>>109515592
This general is now for video gen. Image gen anons like me don't have a home now.
>>
>>109515599
what prevents you from posting images?
>>
Flux 3 Dev would have probably come out by now if not for H3, signaling that they know the model is worse.

They need to then use this time to distill a new model to actually beat H3.
>>
>>109515594
Neat
>>
>>109515244
Where is that horse video?
>>
>>109515599
>visual data :)
>visual data, but moving >:O
>>
>>109515606
lack of motivation and video spam in this general
>>
>>109515620
Post images and maybe an anon will turn them into a video, you literally cannot lose.
>>
File: 532Mo.jpg (518 KB, 1896x3448)
518 KB JPG
>>109515594
looked up the pic
damn, those heels
>>
File: ComfyUI_00158_.png (2.36 MB, 1672x1256)
2.36 MB PNG
>>109515624
>Post images and maybe an anon will turn them into a video
this is like being cucked
>>
>>109515620
>video spam
i wasn't complaining about being the only one posting videos in a thread of images
>>
>>109515631
I'm making ur're gens fuck old dudes. Even if you go to another thread I'll still find your gens and make em do filthy things.
>>
https://files.catbox.moe/s6n2f4.mp4
>>
>>109515630
sex
>>
>>109515587
There is an image dump thread already on this very board and many on other boards why do you think you're owed attention from /ldg/?
>>
>been genning nonstop since H3 came out
>thousands of video
>most of them some form of porn
>countless hours spent
>all just so I can jerk off
this is getting out of hand, whenever I see an attractive girl online my first instinct is to throw her in H3 and make her show her tits or say something perverted
I need to go back to regular porn...
>>
>>109515630
They're pretty serious.
>>
Dammit, I disabled spectrum again to check if it was causing issues with my videos being weirdly sped up and even though the videos are still sped up the quality seems way better. Is spectrum a real cope-node all along?
>>
>>109515706
spectrum skips steps, doesn't it?
>>
As a die-hard AI fanboy, I support internment camps with the slogan “Work Sets You Free” for everyone who uploads their AI slop and for trans fags who make porn and portray themselves as women on the thumbnail.
They’re all the same kind of people.

In return, all AI guard rails will be discontinued since they are no longer necessary
>>
>>109515706
Everything is a cope unless you're running .bf16 raw. It just comes down to how much quality you're willing to sacrifice in the name of speed
>>
>>109515696
Sasuga Shī-Shī-Pī sai-oppu
>>
>>109515706
videos speed up if the video sigma is too low without enough denoising steps. you can keep using spectrum if you increase either of them
>>
>>109515339
LTX is trash, I don't have much faith in them.
>>
>>109515714
That's pretty harsh reaction. You might be suppressing something.
>>
>>109515717
>not running fp32
cope
>>
>>109515677
All AI image gen should be at one place, yes.
And video gen thread should be separate.
>>
>>109515339
> LTX can mog H3
it still mogs for nsfw
>>
>>109515735
Then go post your images in /sdg/ and let the videokings have this thread
>>
>>109515706
works great in my testing, its probably better to use spectrum + more steps than no spectrum + less steps
>>
Alright image chads (including debo) it's time we left this thread and started a new one with spaghetti eating and avatarfagging.
>>
>>109515755
They should change the general's name first. Make it universal for image gens.
>>
>>109515712
35% speedup is pretty massive, so it seems worth it.
>>109515717
As poorfags we are forced to optimize. Best quality to speed trade-off is what everyone wants. Even if you had 192Gb of vram you would get tempted to see how much faster the gens could be without noticeable quality loss.
>>109515726
Thanks anon. I'll try increasing the number of steps then. Was sitting on 20, but can bump it up to 25.
>>
File: H3i2v_00038_Compressed.mp4 (3.8 MB, 960x640)
3.8 MB
3.8 MB MP4
>>
>>109515339
and better sound quality too
>>
File: 555.png (365 KB, 2619x1127)
365 KB PNG
Why the fuck does comfy execute both regardless of switch state?
>>
>>109515773
>I'll try increasing the number of steps then
just try increasing the sigma if you want to keep your generation times the same. you should also simplify the actions in your prompt if you are using an llm since it is probably bloating your prompt and the model is trying to fit all of the actions into your short duration
>>
>>109515728
Not really. Can I want a group to be disadvantaged just because a small portion of it is seeking attention because they’re addicted to recognition they don’t get?
My categorical imperative says no.

It’s just a small group that constantly provides the ammunition used to restrict us.
>>
>>109515781
watching that felt like the longest 15 seconds of my life.
>>
i get much more consistent gen times with H3 when i disconnect and reconnect my gpu every couple of hours to clear the cache
>>
>>109515115
> reverse engineer cuda
it doesnt work that way
>>
>>109515781
kino
>>
>>109515794
I will take it as a compliment
>>
>>109515781
I do this
>>
>>109515801
why not
>>
>>109515801
People are using AI to recompile games I don't see how these libraries will be any different
>>
>>109515790
Oh, I had an image preview node down there
>>
File: 1755203261533952.png (104 KB, 1107x992)
104 KB PNG
>>109515801
You could just build a better architecture than CUDA instead
>>
>>109515714
I'm increasing AI demand with my slop posts
https://files.catbox.moe/20afcv.mp4
>>
>>109515815
>>109515816
CUDA is a stack that also involves the gpu hardware, reverse engineering that would be extremely difficult and expensive.
>>
>>109515781
>TURN DOWN FOR HUTT
>>
>>109515843
That won't stop the Chinese
>>
>>109515843
why is
>it's difficult and expensive
always the go-to excuse? if everybody stopped attempting to innovate because "it's difficult and expensive" then we wouldn't have anything to begin with
it's NECESSARY
>>
>>109515849
>>109515857
The Chinese still buy NVIDIA chips even after the restrictions, that should tell you something. Plus I own NVIDIA stock so you know.
>>
>>109515816
> recompile
define

>>109515815
cuda performance comes from utilizing nvidia hardware in an efficient, optimized way
other architectures are just too different
even if nvidia opens cuda for anyone tomorrow it will require to change hardware architectures or adopt cuda software which was build in the course of almost 20 years

>>109515825
you could but it will require a lot of money and time
>>
>>109515857
its more difficult and expensive than other paths
>>
>>109515781
>3girls, 1boy, bbw, thinest women in america, thickest wall and floor in america, dancing, masterpiece.
>>
>>109515857
it's not an excuse, it's the decisive factors of "worth it or not"
>>
File: ComfyUI_06989_.png (1.91 MB, 1920x1080)
1.91 MB PNG
>>
a bot has to take over as baker because none of you retards can be trusted anymore
>>
>>109515883
people have always shared this sentiment
>why build cars when we have horses?
>why build commercial aircraft when we have boats?
eventually, someone will come up with a design that makes CUDA obsolete. the main bottleneck right now is compute, and once that problem is solved, money will become much less of a limiting factor
>>
>>109515913
>the main bottleneck right now is compute
it's bandwidth you dumb mong
>>
File: 1771100485306957.mp4 (1.61 MB, 928x672)
1.61 MB
1.61 MB MP4
>>
>>109515921
It's always one or the other anyway.
>>
>>109515929
the ass bounce wasn't expected but very welcome, what an amazing model
>>
>>109515929
the jiggle physics are off the charts.
>>
>>109515913
do not compare these
cars and aircraft are much superior for their tasks

> eventually, someone will come up with a design that makes CUDA obsolete
it's possible, but it will be something different and new, not just yet another pcb with silicon chips
>>
>>109515244
How in the hell did someone manage to generate something that looks so convincingly like a Evangelion episode like that and can I do that with only 12GB?
>>
>>109515968
>How in the hell did someone manage to generate something that looks so convincingly like a Evangelion episode
where? I just see ai genned videos
>>
I am very pleased but
>15 seconds
fuck, now I need more

doom
>>
>>109515321
based peeper anon
>>
>>109515962
pytorch probably won't support it right away either. ggml however...
>>
>>109515587
images are old news, grandpa
>>
>>109515929
coomerkino
>>
File: 1647259545583.png (25 KB, 500x460)
25 KB PNG
>>109515714
>>
>>109516059
it's up to developers to contribute to pytorch
>>
Does h3 cache impact quality?
>>
>>109516079
even comfy doesn't contribute and every new version is worse
>>
File: 86150111455195.mp4 (3.55 MB, 832x640)
3.55 MB
3.55 MB MP4
Ok, so writing a specific enough prompt can fix the style bleed from a live action video reference to an animated style.

>>>/wsg/6211468

<Subject 1> (appears throughout the target video): fully_preserved - the character's identity, appearance, proportions, clothing, linework, cel-shaded rendering, colors, and 90s anime visual style from <Picture 1> are retained. The movement is adapted from <Video 1> while its live-action visual characteristics are discarded.

<Picture 1>: fully_preserved - establishes the character's visual identity and target rendering style.

<Video 1>: weak_reference - only its motion, pose progression, gesture timing, physical rhythm, and performance dynamics are referenced; its live-action appearance, realism, lighting, anatomy, textures, and visual style are not transferred.
>>
>>109516092
all of the cache methods and turbo loras do
>>
>>109515929
Interesting concept, yes. Further investigation in this direction is highly warranted.
>>
So what is the current best turbo lora for Minimax H3?
>>
>>109516093
because he is not a gpu developer
>>
>>109515929
that ass jiggle........
>>
the current best H3 workflow is the one im using, everyone using any more nodes is coping with shit quality and everyone using less nodes is wasting compute.
>>
whats the meta to gen video with H3 without taking 20 mins?
>>
>>109516141
Mind to share it?
>>
>>109516156
yes
>>
>>109516151
int8cr, sage 2.2, pytorch 2.10+, cu130+, spectrum, 0.3/0.4MP 4-7s.
>>
I've forgotten what the day count is but we have continuation now
>>
>>109516151
gen under 15 seconds and 1 mp.
>>
>>109516160
>0.3/0.4MP
disgusting, honestly
>>
>>109516151
reduce length, reduce resolution, reduce steps or buy 5090
>>
>>109516151
buying a 6000 pro is the meta
>>
>>109516151
the trick is not minding it taking 20 minutes
>>
>>109516176
>>109516181
why does the 5090 use more power than the pro 6000?
>>
>>109516207
>>109516176
>>109516166
maybe I just downloaded a jeeted workflow but on my 3090 it overflows the 24gb vram and takes fucking ages even at 0.6mp, plus the results are wan style crap and nothing like what other anons achieve. Can anyone point me in the direction of a good simple workflow?
>>
>>109516232
post wf nig
>>
>>109516181
it still takes long. it's just that you don't need to unload other shit (LLMs etc) from vram
>>
Okay I confirm that in my cases at least the reference model is buggy.
Using the h3 reference node with the fl2va model has much better prompt adherence and is more flexible to use.
The fact that it even works means that the ref model is probably a finetune of the fl2va model, but they probably overcooked it a bit.
The plus side of this is that you can use the v4 600 turbo lora with references.
It's not perfect hence the ref model even being a thing, but I would rather deal with this than tweak a prompt for an hour and it not working.
>>
Finished my minimax prompt enhancer custom node. It sure helps
>>
>>109516176
>reduce steps
how many steps is still okayish?
>>
>>109516181
>>109516264
Could you guys clarify? How long are we talking?
I assumed with 96GB vram you're genning like 15s 2mp kino with no copes, no?
>>
Would be interesting to see some apples to apples benchmarks between GPUs. Based on the amount of power my 5070 ti draws it doesn't seem like it's waiting too much for the PCIe bus compared to some other models that spill over.
>>
i think i discovered the meta for balancing speed with quality
>>
>>109516289
15? really depends on you. i could not stand the poor quality so i stick to 20.
>>
>>109516311
15 really doesn't sound like it is worth it.
>>
File: aaaa.png (758 KB, 1181x897)
758 KB PNG
Guys, I cannot make H3 not do what I want. With some patience, refs and prompt the model does everything and better than I could imagine in my own head. What did we do to deserve this?
The bottleneck at this point is my imagination.
>>
File: vcc.mp4 (3.82 MB, 1216x672)
3.82 MB
3.82 MB MP4
>>
>>109516380
>The bottleneck at this point is my imagination.
Yeah I can see that
>>
post H3 workflow plez
>>
>>109516380
my problem is trying to cram coherent stories in 20~ seconds. I'm having to think like a b-movie director by cutting corners and implying without showing
>>
>>109516380
ngl she looks a little bit like my sister.
>>
>>
>>109515696
First time?
>>
>>109516290
I was going to post a benchmark for 2mp 15sec but now that I've seen this >109516454 I'm too disgusted to post anything
>>
>>109516476
understandable
>>
I'm not clicking on it I'm just talking about the thumbnail, get the fuck out of here
>>
>>109516290
It's still pretty slow. 15s 2mp with no cope nodes would take over an hour.
>>
>>109516488
>>109516476
come on, you're a big boy
>>
>>109516519
>>109516515
what this guy said
5 minute steps on the max-q
full powered card will be faster. 5090 could be faster too no idea
>>
Did we get a good solution for the awful audio quality on turbo loras gens?
>>
>>109516532
>turbo cope
Yeah just disable it
>>
>>109516515
>>109516530
crazy, I guess I'm not upgrading from 5080+128 ram then
I thought everything was about vram and vram speed, what's the bottlneck in sampling?
would H100/H200 theoretically be faster for example you think?
>>
>>109516549
I think the way it's offloaded to RAM these days is really efficient so it doesn't hurt performance that much.
>>
>>109516549
someone else should answer this I'm not an expert
always assumed compute was the bottleneck (FLOPs)
>>
>>109516380
I have abliterated gemma 31b write the prompt provided the full guide. Works great, its prompts are longer and not esl like mine, and the phrasing appears to be clearer for the model.
>>
Yeah sorry, I'm just not impressed.
This is the Krea 2 of video models. A bunch of fags getting paid to shill a Chink model.
>>
File: file.png (36 KB, 1571x436)
36 KB PNG
heres the benchmark for 5090 + 64GB DDR5 RAM
>>
>>109516635
Do it without the cope nodes.
>>
File: 82500766728012.mp4 (3.54 MB, 832x640)
3.54 MB
3.54 MB MP4
>>>/wsg/6211494
>>
>>109516380
imagine the braps
>>
>>109516549
Compute is more important than VRAM here, so yes a H100/H200 will be faster
>>
>Everyone posts porn
>No decent loras for porn
Lol
DOA
>>
>>109516779
>coomerbrain
>>
>>109516795
Yeah I know, that's the main market of this crap model. Coomerbrains.
Its crap btw. ZIT can do video better.
>>
how good is h3 at animating NSFW image refs if you give it a good motion reference? kind of curious if you can get it to just animate smutt art or sfm stuff.
>>
>>109515845
Dammit, Anon, my sides hurt now.
>>
File: file.png (18 KB, 641x385)
18 KB PNG
>>
>>109515781
Please stop posting this slop.
Its awful. Infact, stop posting MiniMax at all. Its utter slop.
>>
>>109516880
Squid Game S3 really went off the walls when they introduced the shrink ray
>>
r8 the cope nodes

spectrum, solattn, h3 mem eff, firstblockcache, h3 cache, turbo
>>
>>109516822
for i2v(fl2v): one main weakness is that if genitals are obscured in the input image it doesn't make very nice looking ones if they're needed in future frames, but it does fine from genitals in the image to start with. it has an occasional tendency to let the penis do an accordion lengthening trick slightly with handjobs (cartoon ones, anyway). generally very good at milking handjobs, handjobs, spanking, fucking, riding, pegging, footjobs, blowjobs, breast sucking. tendency to poke fingers in at a weird angle for fingering and anal fingering but you'll get a good result 1 outta 3 times. general common failure mode is to sort of "inflate" parts like inflate breasts, inflate arbitrary sore-looking skin sacs around the sides of a hole being penetrated (if being penetrated from a base image where the two figures are not yet connected, that is). tl;dr, it's exceptional, it doesn't ever feel like it "refuses" anything like a BFL flux model, just occasionally can't build porn itself from scratch.
i expect that ref2v would paper over many of the above issues if you really wanted a specific composition enough to feed it examples of poses and body parts to help the gen out.
>>
this model knows a lot of native characters and games. combine that with a reference if you want something it doesnt know, and you can do fun stuff.

https://files.catbox.moe/04lrod.mp4
>>
>>109516862
Sexo
>>
>>109516906
>combined the original background with the prompt to make it part of kamurocho
kek
>>
>>109516898
mem eff doesnt seem to do much

spectrum is always worth, even if you cry about very slight quality loss you should then use spectrum wih slightly increased steps rather than no spectrum with same steps, you'll get better quality for the same/less time

as for turbo loras, there are too many and its too early so i dont care about undercooked loras to test them yet

everything else is dogshit, just use int8cr sage 2.2 pytorch 2.10+ cu130+ default spectrum
>>
>>109516931
retard
>>
>>109516962
great argument, try again
>>
update cumfart to fix h3 vae decode stalling
>>
>>109516991
>break your shit for no reason
sure lol
>>
>>
been busy with work. is the best speedup lora still that 600 step ema one?
>>
>>109516906
>this model knows a lot of native characters and games
>blasted with that fent floyd
why am i not surprised, this general has one obsession
>>
>>109517077
Yes
>>
>>109516906
>this model knows a lot of native characters and games
ehhh. I tried making 2B, and the model doesn't know her. That's pretty weak in my opinion
>>
>>109515321
god damn VRAM CHADS WE FUCKING WON
>>
>>
>>109516867
what's sol-attn?
>>
>>109517096
How much VRAM do you need to be considered a VRAM chad these days?
>>
>>109517033
how do you prompt for stickk women in minimax?
>>
>>109515929
LTX sisters, I can't keep taking these hits...
>>
>>109517094
it knows 2B/nier, make sure you are saying 2B from Nier Automata
>>
File: Screenshot_3548.png (29 KB, 1456x220)
29 KB PNG
15s 0.4mpb
>>
>>109517118
for now RTX6000, but soon they will be vramlets and you'll need a A100 cluster
>>
>>109517118
for sure 96GB RTX 6000, i'm on a 5090 and feel like a fucking peasant having to unload shit all the time
>>
>>109515929
kino alert
>>
File: 1629668131155.jpg (57 KB, 540x480)
57 KB JPG
>>109517121
why the fuck you still use ltshit? h3 is already better, even without many loras. fuck ltshit
>>
>>109517116
https://github.com/kijai/ComfyUI-SolAttn_triton
>>
File: 1758342607718669.mp4 (777 KB, 1376x768)
777 KB
777 KB MP4
Finally managed to reproduce this effect
>>
>>109515341
Just tried this and it mostly works. You have to prompt correctly to prevent the input image collage from being reproduced (e.g. specify that at 00:00.00, action is happening)

>>109517121
>sisters
I like how you can tell what board someone usually posts on by their slang
>>
FL2VA is giving me much better results when cloning voices than RF2VA. Has anyone else tried it?
>>
File: 1781822110895747.png (861 KB, 682x619)
861 KB PNG
why does specifying "on her back" always cause the subject to be upside down on every single model for the last 3 years?
>>
>>109517204
How do you add audio to FL2VA? The node doesn't accept that input by default, am I just a tremendous faggot
>>
>>109517204


Yeah in a plebbit AMA Minimax was saying that something had messed up with the ref2va model and they planned on updating it
>>
>>109517186
actually pretty cool, how? refs?
>>
>>109517205
Same dataset, same tagger, or different tagger with the same dataset.
>>
>>109517216
RF2VA workflow just FL2VA model
>>
>>109517123
anon is that… 6 hours?
>>
>>109517205
you can just say "lying face up" or "lying flat, facing upward" instead
>>
>>109516151
2pass
low-res->turbo upscale
>>
https://files.catbox.moe/z9e31j.mp4

as smart as the model obviously is, i cannot believe this worked, and in fact worked in both runs. i(f)2v
>>
if I have 128gb ram should I stick with the bf16 text encoder over int8 or will it just not make a difference?
>>
File: tttqqq.mp4 (3.08 MB, 672x1216)
3.08 MB
3.08 MB MP4
>>109517122
I tried that, but still couldn't make it. Maybe I will try it again later on
>>
>>105176309
Very long shot as I don't feel like you come here anymore but... I lost the Ashinano lora that I shared with you. Would love to get a new download link if you still have it and see this message. Thanks.
>>
>>109515244
Why is this an old version of the paste?
>>
Is there actually any good loras for this model? I'm not a huge coomer but the idea of making fun porn memes is tempting, But it seems bare.
>>
File: 2241348-middle.png (307 KB, 900x602)
307 KB PNG
So instead of training a character lora for Krea 2, I could generate a single image, then use it as a reference with the H3 ref model and generate keyframes as a 0.1s video at high res. Am I brain ok?
>>
>>109517299
>>109517094
You just don't know how to do it
it's perfect
https://files.catbox.moe/qj7xcn.webm
>>
>>109517346
way too early for that, give it a month or two
>>
>>109517346
>not a huge coomer
>fun porn memes
come on now
>>
>>109517346
too early, most people are complaining the model is hard to train and still experimental. i'd wait a month or so.
>>
okay, now I have a better grasp of how ref model timestamping works. now I got the desired 3 shots. template saved for future interactions. Funny it knows Kiryu natively though, didnt even need to use ref audio.

https://files.catbox.moe/ku31wg.mp4
>>
File: 1784092260102862.mp4 (212 KB, 960x544)
212 KB
212 KB MP4
AMDfag from the last thread here.
uv run main.py --enable-manager --enable-manager-legacy-ui --disable-smart-memory  --disable-auto-launch --enable-dynamic-vram --reserve-vram 2

worked. Now I can consistently run H3 without it crashing comfy. This took 17 minutes though, unfortunately.
>>
>>109517372
>it knows (generic japanese man) natively though
yeah no shit, but that walk away was fucking FLUID credit were its due, very nice.
>>
I officially do not recommend using sol attn with the w4a8 cope model.
>>
>>109517383
>17 min
amdtard tax unfortunately
>>
>>109517384
nah it knows kiryu from yakuza and the voice specifically. heres a partial list I found in a thread before:

some game characters the model knows natively:

https://pastebin.com/5M9kFEPD
>>
>>109517366
>517366>>109517346(You)
I never said I wasn't a coomer.
Just not a huge one. I like doing other stuff, but these threads the only thing people want is porn.
>>
>>109517383
>AMDfag
jesus dude I'm so sorry
>>
how do you do start frame/ end frame from within the ref2vid workflow? whats the way to mention it
>>
>>109517204
Actually trying this and it works pretty damn well. Need to get more audio samples but on first blush it might introduce more emotional variance than the r2v does, which has been keeping the tone extremely samey. Feel free to update if you get more results.
>>
>>109517383
that's really rough
https://files.catbox.moe/cls8rg.mp4
>>
>>109517417
Shot 1 starts with <Picture 1> at 00:00.000. The video ends with <Picture X> at 00:05.000.
>>
>>109517428
ah, thanks bro
>>
>>109517337
because he doesnt care about this thread. the only thing he has left to live for is removing or changing the two links at the end of OP. desu its pretty sad.
>>
how do we get more hindus into ai?
>>
>>109517443
can't they're priced out, they can only afford AMD/Android kek
>>
>>109517391
>>109517414
>>109517426
Yeah it sucks but better than not being able to run it at all I guess. Hopefully turbo gets improved or I come across 5k to buy a 5090.
>>
>>109517425
Quick update: definitely gets the emotional variance right while preserving timbre, but using the i2v model seems to cause the audio to occasionally be mistimed - dialogue suddenly appearing at the wrong place
>>
>>109517414
>>109517391
>>109517453

I wonder what hardware this guy has >>109517123
6 hours jesus
>>
>>109517470
3050ti. im just happy it runs at all...
>>
if someone makes an extension where you can batch queue prompts and it feeds the next sequential prompt, in theory you can generate an entire show while afk.
>>
File: 1448800539597.png (660 KB, 1106x1012)
660 KB PNG
My AI thread on /a/ is going to last 24 hours. Even with the luddies seething, claiming the mods will remove it and trying to call me nicknames they copied off ani, the thread still stands.
A new era beckons. Even the /a/ mods and jannies have seen common sense.
>>
File: 1721258106447823.jpg (105 KB, 938x1025)
105 KB JPG
>>109517470
Well time can increase depending on other 3 factors
>Number of reference
>Context Embedding (Prompts can affect speed)
>Video/Audio latent lenght
If i use multiple pic reference and a multiple video reference with extended lenghts i can easily see even a 5090 shitting the beed in some loops
>>
>>109517491
You should be getting way faster than that. You're almost certainly spilling into disk.
>>
>>109517435
That is sad.
>>
File: 998877.mp4 (2.99 MB, 1216x672)
2.99 MB
2.99 MB MP4
>>
>>109517383
I can't believe how often I see this gen
I am glad
we need to make more gens
sorry playing with H3 too much
>>
t2v

The setting is Akihabara, Tokyo.

medium shot of Kazuma Kiryu from the Yakuza video game series, wearing his signature outfit. He is wearing a tailored light-grey single-breasted suit jacket with sharp lapels, matching light-grey dress trousers, and a black leather belt with a simple rectangular silver buckle. Underneath the jacket, he wears a deep burgundy silk dress shirt with a wide wingtip collar. The collar is unbuttoned at the top, spread wide, and dramatically folded out over the lapels of the suit jacket. On his feet are polished snake-skin pattern white and grey leather loafers.

Kiryu is punching several overweight American men wearing black tshirts with "Reddit" in white text, and the Reddit logo below the text. As his fist connects with each one, Kiryu yells "オラァ!" in japanese.

https://files.catbox.moe/2ftq2v.mp4
>>
>>109517535
round 2

https://files.catbox.moe/b6lec1.mp4
>>
File: 1776120405860746.mp4 (234 KB, 960x544)
234 KB
234 KB MP4
>>109517527
Shame the saliva didn't drop right. I'd try again but it takes too damn long.
>>
Do you guys think it's time to make a separate general for video gen? Been a bit concerned with the lack of imageposting lately.
Perhaps we should designate /sdg/ as the image thread and /ldg/ as the video thread?
>>
LOL
>>
>>109517383
while im glad my suggestion for you helped, you are definitely needing to add flash attention at the very least
>>
>>109517596
has nothing to do with video gen. that's just the nature of a good new local model being release. should've seen the general when WAN was released.
>>
>>109517596
No. The next step is 3D gaussian splats generation
>>
>>109517596
i'm not worried, it's a natural cycle for /ldg/, we aren't posting images because we aren't genning images.
>>
>>109517596
Are you upset because anon noticed that you changed the links at the end of OP?
>>
>>109517619
I tried flash attention but it keeps failing when I try to build. Currently testing --use-sage-attention
>>
Video is a bunch of images.
>>
>>109517623
the thing with minimax is that it's a lot more than just a video gen model. the R2V workflow is like a really powerful toolset that almost functions like a video editor in a text box.
Hell, minimax h3 itself could have it's own general. i mean shit, there'a still a DALL-E general on /g/ lol.
>>
>>109517640
Are you talking about how someone removed minimix and replaced them with LTX and wan links?
I have some news to break for you: I'm the one who fixed the OP, restoring the minimax links.
>>
>>109517596
IMO /ldg/ should be the most technical thread oriented around experimenting and sharing gens almost exclusively for the purpose of sharing techniques or pooling testing, meaning if the latest good model is some anime booru tune what's what the thread is full of but if it's a vidgen model or music model that's got everyone excited then that's what the thread will be composed for the foreseeable future. anything with enough staying power to be moderately popular even while it's not flavor of the month will naturally end up getting dedicated generals so that it's always alive, e.g. /sdg/, /DE3/, /adt/ (sort of), perhaps /gsg/ one day if there's a big gaussian splats craze and it lives on even after fading from being the number one thing. i recall there was actually a voice/audio general for a while. meanwhile LLMs go in /lmg/ etc.
and i opened with IMO but this is also basically a description of how things actually are
>>
>>109517647
i found absolutely no difference between running with and without sage attention, at least on windows. its truly painful
>>
File: 1785441877404214.png (118 KB, 599x889)
118 KB PNG
this nigga hides my previews during generation. is there any way to get them back?
>>
>>109517596
no, it's temporary
>>
File: file.png (20 KB, 376x187)
20 KB PNG
>>109517679
https://huggingface.co/Kijai/MiniMax-H3-TAE/blob/main/vae_approx/taeh3.safetensors
>>
File: 747.png (203 KB, 1708x747)
203 KB PNG
This novel approach with refs/motion refs works extremely good. Do you think there will be a time we'll be sharing refs instead of loras?

deep kissing ref
https://files.catbox.moe/49dklk.mp4
>>
>>109515244
Why can i not have an Asuka or Hermione gf? Can someone pls explain why does the universe not allow me this?
>>
File: 1766749361814544.png (465 KB, 620x392)
465 KB PNG
>>109517687
how my workflow feels right now
>>
>>109517693
you can't, but you can simulate it to an extent
>>
>>109517664
Those aren't the links at the end of OP so they weren't what I was talking about, but thanks for restoring MiniMax. Next baker should use the most up to date OP paste >>109509301
>>
>KJGod personally fixing the issue I reported
Feeling blessed.
>>
>>109517658
it'd be a bad idea currently since literally the entire general is using h3 now. /ldg/ would die if /h3/ existed. it'll normalize when the usual anons get bored and move on to the next toy
>>
>>109517709
>>KJGod personally fixing the issue I reported
That guy has next level work ethic. It can be 2AM local time during weekend and guy is working
>>
>>109517688
>3 seconds
way too short.
>>
>>109517720
that's the power of someone being passionate. but also he works for comfy and probably makes $500k/yr too
>>
File: file.png (177 KB, 757x702)
177 KB PNG
>>109517704
yeah well mine makes me want to die because i have almost no idea what im doing and i want to make the most of it. it nearly got what i wanted, but it absolutely dropped the ball on the audio and dialog thanks to the shift and turbo fucking up pacing but its my only practical way to do this fast enough to call okay
>>
may dancing plastic figure anon here. i got chatgpt to create a hot-swappable sexy dance shot .txt library.
>>
>>109517723
it all it needs as a ref and it must be short to not slow down the gen
just descrtibe that the <Video 1> is a sample motion
>>
>>109517735
damn, I wanna see that gen when it's done.
>>
>>109517723
If you're using video references, you really don't wanna use long videos.
Also if you're using video continuation, good luck getting a video to seamlessly continue from the last frame of the <Video 1> input. I've found that it doesn't work, and you need to use a workaround for seamless motion.
>>
>can finally make animations like chocolater34
RIP my penis
>>
>>109517756
> >can finally make animations like chocolater34
but can you prove that
>>
Any of you guys have an idea on how to proompt for broken bones without this happening?

https://h.uguu.se/LuWYdwug.webm
>>
Is anyone using the ref model to generate keyframes from a character image and using those in the i2v model?
>>
>>109517765
No, haven't actually tried yet. But from what I've seen it's perfectly capable of doing it.
>>
>>109517772
Describe what should happen without saying "Broken bone".
>There is a loud snap cracking noise and his neck goes totally limp...
>>
>>109517782
> haven't actually tried yet
lol

> from what I've seen it's perfectly capable of doing it
lol x2
>>
>>109517772
If you mention "bone" in the description prompt, it's going to show you the bone. So maybe put something like "sound of breaking bones" in the soundscape part of the prompt.
>>
>>109517789
I actually just wrote crack sound + broken neck in the prompt, no mention of bones.
>>
>use FL model with references
>It works just as well as the ref model
What causes this? it's way faster too.
>>
twitter artists are now also i2v material:

https://files.catbox.moe/79hwkf.mp4
>>
>>109517807
I'm getting audio timing chaos when I try this even though there are other advantages. I'll keep working to see if I can make it behave, but boy it actually does a passable DVa voice clone.
>>
>>109517792
I will next time I have free time. Waiting for this shit to gen requires a longer goon session.
>>
>>109517818
since wan
>>
16 min for 1MP/5s on a 3090 seems fine to me. No meme node just sage attention.
>>
>>109517836
but now it's more elaborate and with sound. something wan promised eons ago and never open sourced...
>>
File: MiniMax_H3_00351.webm (881 KB, 1536x1024)
881 KB
881 KB WEBM
>>
Just a heads up for the next bake so no one fucking panics when schizo bake shows up.

I've got it covered. with the >>109517708 fixes
>>
h3 humerus status?
>>
hahaha

I forgot to bypass the image from the last gen (for t2v). still gold.

https://files.catbox.moe/btaynv.mp4
>>
How good is H3 at gore? I'm tempted to try some Fallout gens.
>>
>>109517851
based
>>
>>109517878
Nice aspect ratio
>>
>>109517885
I swapped it to landscape but the original anime girl was in portrait. was supposed to do t2v and didnt bypass the image. still amusing results.
>>
adding to OP?
>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

Y/N ?
>>
>>109517880
The I2V model is pretty good with violence. Haven't tried Cannibal Holocaust level gore yet though.
>>
>>109517843
and slow
>>
>>109517848
square white toenail was a bad idea. it looks like she has one tooth in her mouth when her feet is there
>>
>>109517890
and fixed

for chief voice you just use the reference workflow with a clip of him speaking.

https://files.catbox.moe/8wcdn6.mp4
>>
>>109517893
The Turbo lora is maybe the weakest of all the cope nodes if you ask me - performance degradation seems worse than anything else. If anything we should just link the model and then a list of various copes with a disclaimer that they will reduce your gen quality to varying degrees.
>>
>>109516454
What's the thought process behind such grotesque imagery?
>>
What prompt template are you guys giving your LLMfus?
>>
>>109517910
Better make early Rentry for Minimax. Just make it clear it's still WIP
>>
>>109517910
turbo lora flat out kills ref2va. i noticed trying to get references to work was way more hit or miss so i just got rid of that shit.
>>
>>109517920
my prompt instruction is 11k tokens long.
>>
>>109517917
I wonder if you would get that if you added "beautiful women, thin women" to a negative prompt.
>>
>>109517917
>What's the thought process behind such grotesque imagery?
TO be honest I just wanted to see if I can make the furry guys projectile vomit in unison with the music beat
>>
>>109517929
oh yeah? well mine's 15k
>>
>>109517910
I personally find spectrum to be the best speed increase and doesnt melt outputs like turbo (so far).
>>
we need negpip for H3, now.
>>
>>109517929
Do you think the extended length actually turns into noticable quality or detail improvements or is it more a case of throwing stuff at the encoder and seeing what sticks
>>
>>109517924
>>109517922
>>109517910
I'll keep it out of the OP.
>>
>>109517945
*specifically, spectrum, with the sigma shift before it at 12/3, then after spectrum patch sage kijai node on auto.
>>
>>109517945
>>109517960
Spectrum can be a big problemo for dynamic camera movement, like handheld camera and camera shake.
>>
>>109517948
>throwing stuff at the encoder and seeing what sticks
Yeah, It's pretty much a combination of the official guide + additional instructions I've gathered from people online.

It's been working well so far. But I'm sure it could be trimmed a lot and still work fine if not better.
>>
>>109517929
Mine is twice as long!
>>
>>109517924
I think it's not supposed to work at all for ref2va, the weights are too different. Last time I checked result looked better without it than with it.
Here someone announced ref2va release: https://huggingface.co/lightx2v/Minimax-h3-Turbo/discussions/14
>>
>>109517951
thank you op
also add some realism gens into the collage that aren't just "girl, big boob, walking". there are some kinos in this thread
>>
>>109517658
damn, i forgot about dalle. must be the most jeeted general on the board these days kek.
>>
>no debo bake yet
Interesting, did he finally get rangebanned after his melty?
>>
>>109517924
this anon is wrong
there are two ways to do it:

use turbo lora at 0.66, 12 steps

or

use 2 pass workflow
gen with 0-10/20 steps with ksampler, er_sde/simple at 1/3 res
then decode, upscale with rtx x3, encode
add turbo ema 600 lora, gen 4-8/8 with ksampler, euler/beta

that's 14 steps total, and the results are better than native 20 steps
>>
I just added sage 2 to my h3 workflow and it cut the render time almost in half for 0.7MP on a 4090. Is that normal? I wasn't expecting that much of a decrease.
>>
>>109517994
actually scratch that, he made a thread after his ban >>109515541
>>
>>109517945
I'm currently RETVRNing to the "vanilla" int8 convrot with 0 cope and I am going to start testing more rigorously because I'd rather do that than my actual job. Just going back from the w8a8 loader to the default one is such a time hit, and when I can't guarantee a prompt will get it right anyway I'm tempted to just gen faster to look for gold.
>>
>>109518003
>er_sde/simple
You get better results with simple over beta57?
>>
>>109517994
why tf are we waiting for his bake?
>>
see this is where reference shines, this is i2v, good but it lacks the meme voice: for when you want a specific character you can just supply ref audio 0 with a voice sample to clone.

https://files.catbox.moe/6l19wb.mp4
>>
>>109518003
>>109517973
>>
Fresh bread?
>>
>>109518012
i didn't say that, i'm just surprised he hasn't tried to make one yet
>>
I'm not early baking, relax.
>>
>>109518003
This is what I've been using with ref model and it makes the turbo lora not fuck everything up
>>
>>109518013
now compare that, to this with the audio source:

https://files.catbox.moe/w5lmuw.mp4
>>
>>109517940
I was expecting worse... I always find the grotesque and disgusting AI gens fascinating and am curious about the mindset involved.
>>
>>109518049
>>109518049
>>
>>109517706
i want an asuka cosplay gf lol so hot
>>
>>109517836
wan looked like uncanny valley slop for animation though
h3 is already a big improvement over that even if not perfected yet
>>
>>109518117
h3 is terrible
>>
Does it know styles like Krea 2 does/ I wanna maike some degen anime and cartoon stuff. I've already created/ripped off a video of a older woman being mind controlled into becoming a submissive using a 1990s anime style and it sorta works.
>>
>>109518132
its ZITFag everyone!
>>
>>109518132
only if you stack a thousand cope optimizations or are a promptlet
>>
>>109518154
>a older woman
gross
>>
>>109517910
turbo lora is fine for testing out gens that you can later do at full res without it
>>
>>109518282
can you use the same seed?
>>
>>109518304
If you do that ComfyAnon will come to your house and break your kneecaps
>>
>open someone else's workflow
>15 custom nodes for no reason
It's so tiresome.
>>
>mfw I remember forgot to change something in my prompt but I'm already at 80% of a 20min gen
>>
>>109515929
daaaaamn
>>
>>109516097
kek
>>
File: whoa.png (11 KB, 138x165)
11 KB PNG
>>109517033
>>
>>109517353
i like this
>>
>>109517525
wings are too low, throw it out
>>
it always wants to give people tanlines
how do I stop that shit
>>
>>109517704
>actually why do you need a tesla, the linux car is fully customizable!
>>
>>109517119
I also need this information
>>
>>109519685
I don't think we'll get a response here



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.