[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109476286

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
great, another troll thread
>>
does turbo load before or after all the sageattention and cache stuff?
>>
I hope wan go bankrupt cause they lied about wan 2.5 open source ages ago

their model is pointless now. thanks minimax.
>>
>>109477686
Is it normal to be continuously disconnected from my home internet when generating?
>>
Is video continuation really broken? anyone else tried it?
>>
>>109477700
you should be containerizing cumfart too
>>
Guys I need more VRAM, any tips?
>>
>>109477699
this
>>
File: SwSh_Lass.webm (1.66 MB, 896x576)
1.66 MB
1.66 MB WEBM
https://d.uguu.se/rzeJNidx.webm
>>
hey did anyone get owned by the npm SC attack through comfy?
>>
>>109477697
>the bread with valid collages is troll but the one with a goofy gif of AI will smith isn't
you're a funny guy aren't you?
>>
File: 1784474488539036.mp4 (1.45 MB, 960x544)
1.45 MB
1.45 MB MP4
>>109477697
>>>/wsg/6209100
>>
I'm using the pruned h3 model now with the turbo loras, but all 3 of the turbo loras are fucking up the audio and introducing the classic floaty shit all over the video.

Is there a fix?
>>
File: MiniMax_H3_00217_R(1).mp4 (3.39 MB, 1088x1920)
3.39 MB
3.39 MB MP4
>5090, 2mp, 8s length, sage att, h3 cache, no turbo lora, 22 mins gen time
genning on 2 megapixel does fix the small faces and grainy artifacts.
>>
>>109477737
no
>>
>>109477744
so you are the /wsg/ troll?
>>
File: 1784931815639948.jpg (1.5 MB, 1248x1824)
1.5 MB JPG
The localgen leap between now and a year ago is fucking insane, year ago we still had to cope with Chroma being a slow piece of ass for any sort of creative prompts, the videogen was painful and anime was still stuck in sdxl hell.
>>
>>109477723
nice
>>109477744
kek
>>
>>109477744
lmao, nice
>>
>>109477758
how would you know? It was to steal API keys. It would affect partner nodes and comfycloud
>>
>>109477699
i still don't understand why wan went completely closed source. they could've just followed the current meta, open source the weaker model but keep the best stuff behind an api
>>
any good h3 loras yet?
>>
God prompting is a pain, I miss the days of wildcard batch-genning 1girls
>>
>>109477768
cause i use wan2gp
>>
>>109477744
cringe
>>
>>109477744
could use more work
good idea
>>
the same fagging is ridiculous over gay thread drama
>>
>>109477744
KEEEEEEEEEEEK
>>
>>109477686
why do bakers always use their own gens for the majority of the collage?
>>
File: it's gonna be useful.png (115 KB, 375x196)
115 KB PNG
>>109477744
absolute cinema
>>
>>109477798
they are narcissistic faggots that samefag like the flamewar instegator
>>
>>109477801
it's not that great CJ. calm down with the samefagging
>>
>>109477764
Sometimes it feels like we're hanging on a thread though.
>>
>>109477798
>>109477804
how do you know the baker's gens? are you a 4chan staff to know that information?
>>
File: 68c.jpg (172 KB, 1106x1012)
172 KB JPG
i unironically nearly pooped myself cause my generation was decoding and i wanted to see it before i ran to the bathroom
>>
>>109477809
>it's not that great
oh I'd say it's pretty great
>>
>search "jav minimum" on Google Image
lol apparently this is a thing in Japan
>>
>>109477813
Surely it's all just automated. I mean no one would pick variations of the same prompt for the collage. No one could be that lame, right?
>>
>>>/wsg/6209107
><Audio 1> is the voice-timbre reference for <Subject 1>
This is absolutely wild kek, is there anything this model can't do
>>
>>109477812
literally the labs and cumfart selling out more and more. they don't even make open source models anymore they are (((open))) commercial licences that keep creeping in more usage restrictions and fees
>>
>>109477827
when was the last time a truly open model was released though?
>>
>>109477824
>No one could be that lame
you can be lamer, you can go for no collage and a will smith gif for example
>>
>>109477831
that Japanese wan fine-tune or I guess Kandinsky is the only one I can think of as a foundation model that's a2
>>
messed up my wf or comfyui
it is taking double the time to gen a video with minimax h3!
tried installing sage 2.2, reverted back to sage 1 but comfy is performing worse now!

need maybe a better workflow for ref! anyone care to halp please??
>>
for some reason I can't get I2V to gen dialogue, even following the formatting on the "official" guide.
I have no problems getting characters to talk using t2v. Am I just being unlucky, retarded or both?
>>
>>109477833
he literally had to make a cope webm over will smith and two anons that are less annoying than him
>>
https://files.catbox.moe/e266l4.mp4

Kek.
>>
>>109477753
Have you tried rtx upscale at 1mp?
>>
>>109477842
that's a man
>>
>>109477833
spaghetti kino is based
>>
>>109477841
no one is more annoying than you
>>
Anyone tried this? https://huggingface.co/Abiray/Minimax-H3-nvfp4-INT4-INT8-Convrot/blob/main/MiniMax_H3_Ref2VA_pruned_mixed_int4_int8_convrot.safetensors

Saves about 5GB in model size, worth?
Same question for the text encoder, qwen3vl. Anyone tried the Q2 quant? It saves 7GB.
>>
>>109477852
must be pretty annoying getting called out like the little bitch you are. you touched catjak lolcow poo and now you are a lolcow yourself
>>
they know too much

https://files.catbox.moe/g6qrky.mp4
>>
make some more will smith spaghetti videos for the next bake
>>
File: choo.mp4 (3.72 MB, 1376x768)
3.72 MB
3.72 MB MP4
>>109477764
i wish new models had chroma data again but that certainly is because things rapidly got better
>>
https://huggingface.co/Abiray/MiniMax-H3-Turbo-Lora-ComfyUI
>The initial release of this MiniMax Turbo LoRA required a standalone custom Python script because standard ComfyUI samplers struggled to handle the dual-schedule (video + audio) mechanics without blowing out the audio track.

>This specific .safetensors file has been structurally pruned and reformatted so it works seamlessly inside standard ComfyUI workflows. You can load it using a standard Load LoRA node. It is optimized to run alongside the pruned or curve-form MiniMax-H3 checkpoints designed for ComfyUI.
>>
>I had to make an entire different turbo because cumfart is spaghetti code
stop worshipping dogshit projects
>>
>>109477844
this?
>>
>>109477842
that's a boy
>>
yo they did some good shit with comfy
changing loras doesnt reload the model
inb4 its always been like this and something simply was fucked on my end
>>
Outputs with the turbo lora are much worse than the optimization node stack for me. Even at 10 steps
>>
File: 00067-46.png (2.71 MB, 1728x1344)
2.71 MB PNG
Ran into Sysram OOM... video too big to combine with ComfyUI. Anyone else autistic enough to link shots together?

https://streamable.com/f3ofca
>>
>>109477878
his video showcase on the model card is terrible, the sound is still broken, I'll just wait for the lora to be finished, as simple as that
>>
>>109477933
the problem with doing this is the obvious framerate differences
>>
File: MiniMax_H3_00449_.mp4 (2.53 MB, 928x928)
2.53 MB
2.53 MB MP4
>>109477817
Understandable desu.
>>
How does H3 not completely invalidate the business model of Onlyfans or the porn industry in general? If someone has even one public picture I can gen a high quality video of them getting railed
>>
>>109477953
I see. That is another thing to pay attention to.
>>
>>109477971
It will eventually, but right now it's too inaccessible for the average normie.

>>109477933
Can't you just have claude or chat gpt write a script to join them outside of comfy?
>>
>>109477983
that was the nice thing about ltx. you could load a series of frames at the start of the clip in order to do a proper continuation. i wonder if that is technically possible with h3
>>
File: MiniMax_H3_00034.mp4 (2.32 MB, 1296x720)
2.32 MB
2.32 MB MP4
>>109477860
I'm using a Q2 quant for the text encoder, but I tried the base Int4 H3 and it was garbage. Using Int8 atm and it seems fine.
>>
ty google ai search I learned something about the reference model and voices.

When you are NOT using an <Audio 1> slot, the model relies 100% on its
internal 32B text encoder database to generate iconic voices.

To trigger these voices flawlessly and prevent the AI from guessing,
you must use direct Language/Character Conditioning Tags (<d> tags).

THE SYNTAX FORMULA:
<d>[Language, Character Name voice from Name of Show, descriptive traits] "Your dialogue here." </d>
>>
>>109477825
Why are APIs so expensive when this 33b model can do all this shit on consoomer hardware in reasonable times?
>>
>>109477998
why won't you use int8 for both the TE and the model? both can't fit in my gpu but it's ok it's automatically offloading, at least the quality is here
>>
>>109477971
Because it's still hard and very time consuming and expensive
>>
File: Krea2_turbo_01061_.png (1.04 MB, 1024x1024)
1.04 MB PNG
Why does beta scheduler straight-up erase steps?
I made my own custom beta scheduler variant and noticed that the alpha and beta values that deviate aggressively from the "simple" no-op baseline (1.0) result in less and less steps.
THIS IS NOT A BUG OF MY CUSTOM NODE. It is present even on the Comfy's default BasicScheduler implementation with beta selected, but since these use more tame values (paper's 0.6 default), you need to crank up step count to see the effect. I don't know the absolute minimum but I can say it is present >666.
Shift seems to have no effect on it, both beta and alpha cause the bug, both for values <1.0 and above >1.0 (non-standard use, but still)
No other scheduler seems to behave this way, EXCEPT SigmoidOffset, which is a furry meme so yeah not sure how much that matters.
If I wasn't a lazy piece of shit I would consider opening an issue. Still, I am very confused.
>>
>>109478008
>33b model
kijai pruned it to 20b without any quality loss, minimax is genuinely smaller than ltx 2.3 while being 10x better
>>
>>109477996
Pretty sure it has support for that in the reference model
>>
>>109477867
He's right tho
>>
>>109477753
dam thats still 42+min on my 4090, maybe due to 20 steps, guess I'll wait a week for 2mp kinos
>>
>>109477933
Probably would be better to gen a bunch of separate videos in the different scenes that manually edit them together. More effort though but would look better and feel more like a real music video.
>>
>>109478021
i don't think so. the closest thing would be to prompt a "direct continuation" of a reference video, but that isn't a real continuation
>>
File: MiniMax_H3_00036.mp4 (2.34 MB, 1296x720)
2.34 MB
2.34 MB MP4
>>109478011
because I have a 3060 12gb and 24gb of ram and I'm trying to limit pagefile/ssd thrashing.
Already pushing the limit with 5s at 0.4mp
>>
>>109477873
This
>>
>>109478008
>Why are APIs so expensive when this 33b model can do all this shit on consoomer hardware in reasonable times?
API had a monopoly on video genning, so they could pretty much charge whatever they wanted. thanks to H3, that has now changed
>>
>>109477873
>>109478039

>>109477809
>calm down with the samefagging
>>
>>109477933
My prompts aren't good enough yet, but I've saved what you've shown of your workflow for later use.
>>
File: MiniMax_H3_00037.mp4 (2.26 MB, 1296x720)
2.26 MB
2.26 MB MP4
>Flappy metal doors
lol
>>
>>109478001
Do you not read docs?
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
>>
>>109477890
yeah
>>
are sparks like GB10 good for H3?
I know people stack them for LLMs and get good performance
can you do that with H3 and use multiple sparks at the same time to speed it up?
>>
>>
prompt sir
>>
File: lyra_00024_.png (1.7 MB, 1536x1536)
1.7 MB PNG
https://d.uguu.se/vmCLhYYq.mp4
>>
>>109478014
Did more testing.
The Comfy default beta scheduler starts missing a step at 139, as you crank the step count up, the number of missing steps increases.
>>
>>109477878
trying this, seems alright so far but needs further testing. i at least like that there's no python script required
>>
>>109477744
GEEEEEG
>>
>>109478086
main.py is a python script
>>
File: 1755072889088213.mp4 (1.65 MB, 960x544)
1.65 MB
1.65 MB MP4
>>109477873
>>109478039
Unable to make your own, Ranjesh?
>>>/wsg/6209130
>>
success! persistence pays off. I had the <d> tags including the other tags. this format works:
>Eric Cartman from South Park waves a police nightstick angrily, pointing directly at the camera while shouting out in an abrasive, squeaky, high-pitched cartoon voice: <d>[English] "respect my authoritah, Floyd!"</d>

https://files.catbox.moe/r4w3i7.mp4
>>
>>109477744
Can you try making a version with subtitles?
>>
>>109478092
I'm talking about the extra python script the original turbo lora requires you stupid idiot. https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora
>>
>>109478094
how to get realistic south park dudes:

<Subject 1> is the character in <Picture 1>.

The setting is the TV show South Park.

The character Eric Cartman from the animated show South Park is dressed as a police officer. From 0 to 3 seconds, he lifts a pink glazed donut to his mouth, takes a large, visible bite, and chews it happily. From 3 to 6 seconds, he swallows the bite, lowers his hand, looks directly at the camera. Suddenly, he transitions into the character Eric Cartman from South Park, gesturing wildly and speaking in an abrasive, squeaky, high-pitched cartoon voice: <d>[English] "ahh, what a peaceful day in the city."</d> he looks to the right and sees <Subject 1> reimagined as a South Park character, featuring a completely flat, round, 2D digital cutout construction paper body wearing a bright winter coat and mittens. However, his head features a shockingly realistic, hyper-detailed human face with 3D depth, natural skin pores, and real cinematic lighting, <Subject 1> is holding a white bag of powder. From 6 to 10 seconds, the character Eric Cartman from South Park waves a police nightstick angrily, pointing directly at the camera while shouting out in an abrasive, squeaky, high-pitched cartoon voice: <d>[English] "respect my authoritah, Floyd!"</d> Eric Cartman hits <Subject 1> several times with a police club, and <Subject 1> falls on the ground.
>>
>>109478080
It misses a step with H3 but doesn't with Krea until much later...
Why does this shit work so weird?
>>
is there anyway to make loras from different models compatible?
>>
why can't h3 make alarm noises?
>>
>>109478099
I am just making fun of you for being a bootlicker
>>
>>109478119
We already know you're an autistic moron, anon.
>>
File: 10678934.mp4 (3.65 MB, 736x1024)
3.65 MB
3.65 MB MP4
>>109477953
>>109477983
Try explicitly defining the frame rate in the text prompt, maybe it'll work (or at least keep it similar enough that it's not noticible).
>>
>>109477770
>i still don't understand why wan went completely closed source. they could've just followed the current meta, open source the weaker model but keep the best stuff behind an api
because everything after 2.2 WAS the "weaker model", it was shit
>>
File: 1984756465479.jpg (310 KB, 852x673)
310 KB JPG
>>109477753
would you share your mayli lora
>>
>>109478115
aWOOOOOOOGAAA
that kind?
>>
>>109477753
this is a retarded fetish
>>
>>109477817
not yheth
>>
>>109478128
no, like modern warning beeping
>>
File: 5431555151.mp4 (1.99 MB, 1216x672)
1.99 MB
1.99 MB MP4
https://files.catbox.moe/czy3u6.mp4
>>
>>109477844
does that work on a 3090?
>>
is 24gb vram + 32gb ram not enough for h3?
>>
>>109478170
you should be fine with a quant
>>
>>109478170
Yes
>>
File: getAjob.webm (2.11 MB, 1143x2048)
2.11 MB
2.11 MB WEBM
>>109478168
It should work with nvidia I used it for this, but I have blackwell, it does eat up my sys ram for some reason
>>
>>109478125
it's not a lora, u just load a reference image
>>
>>109477698
I put mine right before the Sigma shift node at the end. I am pleased.
>>
>>109478179
Which one? I'm using minimax_h3_fl2va_pruned_int8_convrot and minimax_h3_ref2va_pruned_int8_convrot but my system slows to a crawl ones it starts loading the model. I've only gotten it to gen something 3 times but it usually doesn't make it to that point before I have to kill comfy or it crashes.
>>
>>109478115
Are you sure you're not just unable to hear it?
>>
>>109478170
Personally, I have 64GB vram and only 16 of it is used. 0.2mpx videos take 6min to generate. So you should be fine.
>>
>>109478115
I have actually had no luck getting h3 to make any kind of alarms or beeps, have wanted it in a few gens and simply cannot make it work. Post if you figure it out
>>
How does Minmax compared to Google's VEO?
>>
>>109478226
>Google
meme
>>
>>109478214
see >>109478200
I'm also using q4 text encoder but still no luck. Worth mentioning I'm an amdkek but other people seem to have no problem.
>>
>>109478226
veo is dogshit
>>
>>109478226
Only H3 and Seedance 2.5 remain in the ring
>>
>>109477878
tried the turbo with comfy, audio needs more work
https://files.catbox.moe/2im399.mp4
miku vacation if you click at work
>>
>>109478207
>loud alarms beeping for a long time
i have this in my prompt and multiple generations produced nothing
>>109478219
it's very weird. i wonder if they messed up with the dataset captioning
>>
>>109478246
did you try using the chink symbol instead of english?
>>
>>109478200
It first has to load from storage, are you using some trash HDD?
>>
>>109477878
Yeah the audio takes a big hit. It's going to be really challenging to optimize performance for H3 without fucking up the audio.
>>
>>109478231
>amd
I'm sorry, good luck. I've also had some ridiculous problems with comfy that I've solved by talking to chatgpt about it. Maybe it can be of use to you.
>>
>>109478262
good idea. i have no idea which one would be correct, hopefully google translate gives me the right one
>>
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/discussions/5

another turbo lora, these are the people who made the qwen controlnet
>>
File: the-miku-take1.mp4 (2.39 MB, 736x576)
2.39 MB
2.39 MB MP4
https://files.catbox.moe/u5ce7l.mp4
>>
>>109478267
Nope, NVME and DDR5 RAM.

>>109478269
I've tried chatting with a few cloud models but none have been able to help so far. Guess I'll just keep fucking around until something works or I give up. If prices weren't so fucked right now I'd switch to a 5090 in an instant.
>>
Is so over ;__; 40 minutes of gen and still 50% percent
>>
https://rentry.org/ldgcollage_v2
Updated to show thumbnails of external videos (make it easy to remove prompt/seed variations)
>>
>>109478286
cool, thanks for the work anon
>>
For anyone stuck in Windows, are there any suggestions you have to improve memory optimization besides checking inside my anus?
>>
>>109478298
why would you be stuck. it's your own retardation keeping you there
>>
>>109478298
run win debloater and winutil scripts
dynamic pagefile
install nvidia drivers without nvidia app
update pytorch above 2.9, sageattention 2.2 if below blackwell, 3 if blackwell, cu130+
power profile to high performance
buy more ram
>>
>>109478305
why are you agressive cumfart? there's a lot of github complaints about that already, don't put your head in the sand and fix your shit
>>
File: file.png (180 KB, 819x470)
180 KB PNG
H3 turning out to be one of the most difficult models to train is some monkey paw shit
>>
>>109478298
those flags are useful --vram-headroom 1 --disable-pinned-memory
>>
Wouldn't a workflow that first generated a video at 0.1MP or something like that, then used that as a source (scaled to the final resolution) instead of random noise, help speed things up? Or am I missing something here?
>>
>>109478337
>people are still learning that training a guidance distilled model is mission impossible
we knew that since flux.1 dev, in 2024...
>>
yeah i don't know how to make the alarms. i put some translated chink letters into the prompt and it just makes more ambient noises
>>
>>109478341
I'm using this shit in Linux, is useless here?
>>
>>109478354
there's only one way to know
>>
>>109478330 (cont)
check if xmp is on
use igpu for monitors instead of gpu, or at least create a separate firefox profile without hw accel in order to be able to use comfyui without it stealing a lot of gpu resources while genning
>>
>>109478337
should we be worried, or does ostris not know what he's doing?
>>
File: renak1-ns.webm (2.12 MB, 640x992)
2.12 MB
2.12 MB WEBM
https://h.uguu.se/hDAtoKBF.webm
>>
>>109478354
I'm so close to being able to run two instances I might try it
>>
>my dumbass realized days in that i've been running the unpruned 31gb int8 model this whole time and didn't notice until this sleepy morning
haha lol oops but somehow that isn't giving me the performance hit i'd expect, i still managed to beat some redditor's quickest s/it with this card. and less ram even.
>>
I want to generate some video but it is already like 40C in my room... Help me anons...
>>
>>109478337
It's a shame for finetunes but the reference model is so powerful I don't think loras are necessary
>>
Out of all places, on fucking /gif/ do I find an answer as to why my gens keep fucking up when I go higher than 720p.. h3 just seems to not accept high res gens.
>>
>>109478383
>unpruned 31gb int8 model
notice any quality change?
>>
>>109478390
buy a simple fan from ali and point it from the top of your pc to the door so air circulates, buy ac, powerlimit gpu to ~70%, undervolt
>>
>>109478390
just on the air con
>>
>>109478397
i'll find out in 46 minutes, but it's already telling when i couldn't tell the difference at all from all the pruned gens in this thread and online, and the slop i've been making here with the unpruned. i may just be retarded though not ruling that out yet.
>>
File: 556.png (1.57 MB, 832x1216)
1.57 MB PNG
>>
>>109478373
saw someone train it to use a dragon character from a cartoon. at least for the example shown it went relatively well. not sure about overall quality.
>>
File: Krea2_turbo_00584_.jpg (2.67 MB, 2048x3072)
2.67 MB JPG
>>
is the h3 turbo lora worth using or is easy cache still better?
>>
>>109478414
why did you crash out itt again?
>>
>>109478337
Please anon give us a explanation to retards (Me), the model is untrainabled like FLux?
>>
>>109478424
Stop dedicating your life to being a human gnat
>>
>>109478391
they absolutely are necessary if you change the face angle even a tiny bit
>>
>>109478428
i wouldn't jump to conclusions yet, it's just proving to be very difficult because of the guidance distillation
>>
>>109478428
Minimax is like Flux.1 dev, it is a guidance distilled model (it has no CFG), and it's making the model almost impossible to finetune, that's really a shame they never gave us the base model
>>
>>109478414
Damn, if Krea gets a Danbooru tune it'll be almost the best at everything.
>>
>>109478124
i remember when they first showed 2.5 gens, i thought "oh good, they under trained it so it will probably be open weight."
>>
>>109478431
You can give multiple reference images for the same character
>>
>>109478424
What was the crashout?
>>
>>109478447
yes, you can cope, I get that. it's just not the same quality for obvious reasons
>>
>>109478444
animasisters...our response?
>>
File: 1765174142450831.png (1.56 MB, 1920x1088)
1.56 MB PNG
https://d.uguu.se/sCpiCAJN.webm
>>
>>109478438
So is OVER is not? I read the paper and you have right they supposedly will launch the model undistilled later, but for now... this is for me really over.
>>
A problem i dont understand.
Since compiling sageattn 2.2 on linux for cuda 13 on an 8.9 card and turning on the minimax h3 eff sage attention node my gens now "work" instead of core dumping comfy but they are sped up even though inspecting the file it says 25fps. They were all promps that were originally 10s but i ran them down to 5sec to test Sageattn times. I'm testing the same gen now but back at 10s in case there's some "we have to cram it all into 5s!!!" speedup involved.
If this is a natural feature of the model and not a sage problem anons might find it useful to squeeze more fast paced content into a gen if the default parameters are not producing the desired effect. Overask and over time and then set gen duration lower.

Test complete.
It's a feature.
>>
>>109478348
hes not training a regular lora but trying to dedistill the model
>>
>>109478472
hope you're right
>>
>>109478462
oh thats interesting so if it can compress the prompt it must have some idea of the frame budget up front
>>
File: 1774033684077093.png (537 KB, 1079x666)
537 KB PNG
>>109478461
>they supposedly will launch the model undistilled later
wait, really?
>>
>>109478461
they might decide not to now they are trying to take down every nsfw lora or usage
>>
>>109478461
we figured out how to make zit loras when people thought it would be impossible. its not over yet
>>
>>109478472
even simple loras on guidance distilled models are a bitch to train
>>
>>109478482
>now they are trying to take down every nsfw lora or usage
you are gay retarded anal taking faggot fake news, go check civitai.red.
>>
>Triggering PromptExecutor cache reset. Reason: cpu_threshold_exceeded
What is the argument to disable this?
>>
>>109478489
>those loras don't count
>can't train guide free model
>china tricked us
>can't even do nsfw anyway
>loras aren't real training
>show me a finetune
>that finetune doesn't count
>seedance can do nsfw better
>reddit/grok/9monthsago
>REEEEEEEEEEEEEEEEEEEEEEEEEEEEEE
>>
>>109478394
i think minimax is planning on releasing their special upscaler for the model, which is why they didn't train on higher resolutions
>>
>>109477839
nvm, got it
https://files.catbox.moe/mkw55p.mp4
>>
File: 1767692231090352.jpg (8 KB, 225x224)
8 KB JPG
>>109475560
Good morning sirs.
Me just wake up.
Any updates for the last 9 hours ?
>>
>>109478524
bahraim sir released the turbo h3
>>
>>109478524
yeah im still alive
>>
best out of lots of attempts, sexo

Did take 1.5 hours on a 3090 though

https://files.catbox.moe/b1u0xf.mp4
>>
>>109478531
Source ?
>>
>>109478430
projecting
>>
>>109478544
what a complete waste of 1.5 compute hours
bro's got a plume coming out of his dihole.
>>
>>109478544
Hey, haven't seen you since I left /hdg/. How's the 'treon doing?
>>
>>109478524
You wake up at 19? damn
>>
>>109478556
I lost a lot of interested in both that and twitter after the algo started hard hiding explicit nsfw on twitter and all engagement dropped, followers don't even see most of my posts
>>
File: the-miku-take2.mp4 (2.39 MB, 736x576)
2.39 MB
2.39 MB MP4
https://files.catbox.moe/jxqhmr.mp4
Take 2 came out better
>>
What kind of problems should I expect using the Q2 quant of the qwen3-vl text encoder for h3? Will it make prompt adherence worse?
>>
>>109478588
stop using meme quants
>>
File: 2309498385.jpg (43 KB, 1000x591)
43 KB JPG
it makes some good music. i want to recreate them in my music program
>>
>>109478591
>moron can't answer the simple question
>>
File: file.mp4 (867 KB, 864x480)
867 KB
867 KB MP4
>>109477697
>>>/wsg/6209171
>>
File: MiniMax_H3_00265.mp4 (560 KB, 576x384)
560 KB
560 KB MP4
>>
>>109478476
I cant tell from my gens if audio is sped up as well as it was birdsong and a whistle and i dont have a great ear. I'm going to check out a previous gen with voice and music with the duration cut in half from the original.

So, the audio wasn't changed enough to say that the rate doubled, the scene action was physically faster and all the voices had the right pitch near enough and speed, maybe a bit more energy in the delivery.
>>
>>109478298
If you are even remotely serious about AI, you need to atleast dual-boot Linux

ALL AI development both private and academia and 99% of top-level usage (as in gpu farms, enterprise etc) are done on Linux. It's the native platform for all AI.

Which means that it is where you will have the best support and best performance (not hard since Windows is also a bloated mess). If you are just dabbling in AI, keep using Windows.
>>
wan 3.0 is out now and open source
https://xcancel.com/Alibaba_Wan/status/2085339761284104529
>>
>>109478614
I just use wsl :)
>>
>>109478616
>open source
at no point they said it's gonna be open source
>>
>>109478344 (me)
I guess the main problem is that at 0.1MP video/audio quality with MiniMax H3 for some reason is always complete shit.
>>
>>109478588
Just try it. Lower precision doesn't really matter as much for embeddings as it does for long-context text generation, so it should still be good.
But the TE isn't really a bottleneck since it just gets offloaded after processing your prompt.
>>
>>109478616
Kill yourself fucking retarded faggot
>>
File: Rena_Mion_2.webm (1.59 MB, 672x800)
1.59 MB
1.59 MB WEBM
https://h.uguu.se/nOiSblDG.mp4
>>
>>109478616
I see neither open nor source.
>>
>>109478616
the videos look like ass, they can keep that shit for themselves lool
>>
>>109478616
well that looks like shit
>>
this decoder is way too slow considering
>>
File: 1775996993109117.png (551 KB, 728x1077)
551 KB PNG
there is a very positive aspect to the H3: the background and the characters do not change from one view to the next. what an excellent model
>>
It will still produce tokens corresponding to what you said, and I bet the model still mostly "understands" what you say. But the problem is that at that garbage quality, the divergence from the precision diffusion model was the trained on is massive, so the mathematical representations will no longer point to any meaningful direction and you will get errors, deformities and bad results in your gens.
I honestly don't know what use case there is to it, if you are this desperate with the TE, how the fuck are you going to run the diffusion model?
>>
>>109478648
I am too retarded to tag apparently >>109478588
>>
>>109477744
this needs to be posted every thread
>>
What weight/steps are you guys using for the turbo lora? I am getting some artifacts.

https://d.uguu.se/SGxHojEY.mp4
>>
>>109477744
Based post
>>
File: Krea2_turbo_00616_.jpg (2.92 MB, 2048x3072)
2.92 MB JPG
>>109478444
I think you can get a desired style with the correct prompt, so far I haven't needed to use adetailer or high resolution
fix
>>109478605
LMAO
>>
why minimax prompting style is so retarded ?? what happen to natural language style of prompting ??
>>
>>109478616
>Barely H3-tier model
>not open sauce, no DL link
Yep I think it's over for Wan
>>
>>109478160
hummina hummina
>>
>>109478679
the turbo lora sucks. spectrum with 20 steps is much better, unless there's something im missing
>>
>>109478703
>unless there's something im missing
you're not missing anything, the lora is far from being fully trained
>>
>>109478703
turbo lora is still experimental, it's not even finished yet
>>
>>109478703
mines working well enough, but I keep it at 20 steps
sigma shift 12/6
lora 1.3
>>
>>109478689
huh? i'm doing natural language and getting good results. maybe you just suck at writing
>>
will Smith next thread OP?
>>
>>109478714
>but I keep it at 20 steps
Doesn't that entirely negate the point of the lora?
>>
>>109478714
how long do your gens take?

>>109478703
I just tried another gen (with real life stuff instead of animu) and got a good result/speed. Need to test more
>>
>>109478716
<subject> x doing y like <reference2> yadda yadda works better
>>
>>109478685
>>109478624
>>
>>109478720
nope its still faster than control. try it urself
>>109478722
0.5mp ~ 60sec on my 5090
>>
reference model is so fun, I still have to figure out multi image stuff though

the setting is the forest from <Picture 2>, both characters are sitting in the wooden cart from <Picture 2>.

<Picture 1> is sitting in the cart to the left of the blonde nordic character in <Picture 2>. the black man in <Picture 1> says "where is the fent at, nordic brotha?" in a black man's accent. the nordic blonde man in <Picture 2> says "well, there is no skooma here boy, and you cant steal my bike. no bikes in Skyrim, you see." in a swedish accent. A dragon from above shoots fire at the cart.

https://files.catbox.moe/pk2xbh.mp4
>>
>>109478595
>moron
>he says, trying to hit every model with the q2 ggoof retard hammer
honestly even having to ask gguf questions in current year warrants you a new tard helmet.
>>
>>109477686
>migu
:)
>>
>>109478735
0.7* 5sec
>>
File: Krea2_turbo_00618_.jpg (2.79 MB, 2048x3072)
2.79 MB JPG
This model is powerful
Even if the chroma lora doesn't go anywhere it enables a fuckton more freedom without hurting quality
>>
Thank you for your valuable opinion Catjack.
>>
>>109478755
powerful
>mirror
can it also do human reflections!? or are we not there yet
>>
>>109478449
>>109478755
lol Ani is more powerful than you and you squirm like a faggot biting his heels. I'll bet he doesn't even think about you at all.
>>
>>109478765
>I'll bet he doesn't even think about you at all.
but you certainly do mr mysterious ani defender anon (totally not ani btw)
>>
>gen hits you with the vine boom out of nowhere
as if I couldn't get any hornier
>>
File: view.jpg (98 KB, 832x640)
98 KB JPG
>>109478738
I have only used r2v since, and never opened up a i2v.

screencap is from a failed gen but so good.

>>>/wsg/6209178
>>
>>109478769
lol. crash out. easy W. you really are my puppet :)
>>
>>109478774
nice doro
>>
>>109478769
Dude has nothing else going in his life, whomever it is has lost. Notice how he never post his gens?
He's been priced out from the big guy stadium
>>
>>109478775
this really doesn't land when everyone saw you spend your entire christmas holiday chimping out on 4chan LMAO
>>
>>109478776
thank you yes so good
>>
File: file.png (92 KB, 965x691)
92 KB PNG
>>109478689
For technical control. Vibe code something to help with it. Then you can just say
"make bob and vagene jiggle bounce" and have an llm generate your prompt.
>>
>>109478787
I have no idea what you are talking about. did someone else bait you all Christmas?
>>
Can I run MiniMax H3 with 16gb vram and 32gb system ram? I thought it would be smaller but the Q4_K_M.gguf is already 20gb
>>
>>109478771
kino
>>
Can the anti ani schizos please chill for a doy?
>>
Can I run MiniMax H3 with 8gb vram and 640kb system ram?
>>
File: 1776011140213520.webm (2.45 MB, 640x880)
2.45 MB
2.45 MB WEBM
>>109477686

Just a reminder Minimax still cannot do this no matter how hard you tried with your prompt + Guaranteed choppy DBS tier animation.
>R2V
Thats cheating

>>109478811
Int8 Pruned is 19gb
>>
>>109478811
use the int8 pruned model, anything above your vram will use system ram and all together it should be fine, just dont generate at 1.0mp cause thats a lot even for a 5090
>>
>>109478811
i have the same setup as you and it runs fine with the nvfp4 text encoder and the standard int8 convrot diffusion model so yes that'll be fine. be sure to have updated comfyui and cuda and kitchen or whatever, then sageattention 2.2 launch option and optionally 'spectrum' node for another ~50% speedup, which you'll want. and i do have to make sure i'm not running anything else graphically intensive like a videogame or a civit.ai tab, or else it hangs loading the model until i shut them
>>
>>109478826
yeah the chopped animation in h3 is an issue. what are anons doing to prevent that?
and no r2v is not fucking cheating it's the best mode
>>
>>109478790
Theres literally tons of them. Whats the good workflow for it ?
>>
>>109478806
we all believe you mr mysterious ani defender anon :^)
>>
>>109478846
R2V is literally just Video to Video / Inpaint
>>
>>109478826
>Just a reminder Minimax still cannot do this no matter how hard you tried with your prompt
Thank god...that looks fucking horrid
>>
the comfy update made my h3 gens slower by like 25%. what do i do?
>>
>>109478861
You can't do anything if you need to ask what to do.
>>
>>109478826
>>109478846
>chopped animation
That's what anime looks like. I'm not sure what's the complaint here. Every other open source model before this did it wrong and looked like flash tier shit.
>>
>>109478827
>>109478841
thank you
>>
File: ComfyUI_00025NoAudio_.mp4 (2.1 MB, 960x544)
2.1 MB
2.1 MB MP4
>>>/wsg/6209192
Weird warping stuff at the end is annoying
>>
How do you change the color mode for tag autocomplete? This looks like shit
>>
>>109478867
Thats the problem. Anime is just a still image with flapping mouth.
>>
>>109478867
that's a fair point but I think you should be able to prompt the chopped shit out if you wanted though
>>
>>109477686
there should be a minimax node that take last n seconds of video as input so you can extend a video.
>>
there we go, skyrim kino

https://files.catbox.moe/fb9791.mp4
>>
how do i prompt slow zoom out from subject on h3? it keep cutting to the zoomed out scene instantly.
>>
>>109478886
you asked /adt/ last time
>>
>>109478867
This. Most people don't want to gen traditional anime. What they really want is modern, 60 FPS cartoon animation which is basically just stylized 3D. They should stop prompting for anime
>>
>>109478894
>>109478903
I suspect the choppiness is hardlocked with the anime style. You guys might have to wait for some type of lore to come along.
>>
>>109478826
Can you try adding "3d render" and "blender animation" to your prompt to see if it fixes the "choppy"ness? I would but I'm at work.
>>
>>109478922
I did. Still choppy
>>
>>109478886
Mess with its css until you like what you see
>>
File: Krea2_turbo_00628_.jpg (3.02 MB, 2048x3072)
3.02 MB JPG
>>109478886
Use a tool like cursor and ask qwen 3.6 26b to do that
I made modifications to that myself doing that.
>>
>>109478910
use time stamps
>>
>>109478915
I disagree with most people. Most people do want anime, these guys are just got AI brained from seeing shitty AI animation for too long. It's like the WAI sloppers
>>
>>109478880
A daring synthesis
>>
>>109477744
You've done him
>>
>>109478938
i did
>at 00:02.000 the camera zoom out,
>>
What should the system prompt be for Gemma 4 heretic/uncensored to so it knows it's uncensored and can write lewd text?
>>
Why can't I route a ref video into the default ref2vid H3 node? I just tried Load Video as one would assume... does it only take that input from some special node?
>>
>>109478960
did you put a timestamp for the end of the zoom? and try "perspective slowly zooms out"
>>
SInce HANDJOB prompt didnt work for Minimax, what good prompt to do handjob well ?
>>
good goys
>>
>>109478939
>>109478939
>>
The image can have (or has if all your dataset is NSFW) NSFW content. Describe it accurately in correct everyday language using explicit terms when relevant. Do not mince words or talk around these concepts.
Is enough for any halfway decent heretic model.
Add some more descriptions about how long your caption should be, what you want it to mention/what not etc.
>>
>>109478995
kek
>>
>>109478066
good gen
>>
>>109478160
Yesus
>>
>>109478826
the mirko looks far better than whatever this is lol



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.