[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109456130

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Z
https://huggingface.co/Tongyi-MAI/Z-Image

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Wan
https://github.com/Wan-Video/Wan2.2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>>109457662
Good job
>>
>>109457662
Based moving collage.
>>
>>109457662
not LTX 2.3 in the OP ?
>>
Anybody tried under 20 steps yet? What's the verdict here?
>>
>>109457546
I hear nothing
>>
>>109457702
yes, the movements become stiffer as you go up
>>
Blud this spectrum shit is lowkey crazy
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
Circa %30 speed-up. I have only tested a few samples but so far I did not notice any strong deviations or errors in them. I even tried t2v without any reference image and it still spits out close likeness.
I also don't know how it combines with other optimization but you must check it out at minimum. Really missing out if you don't.
>>
>>109457707
UNDER 20
>>
holyfuck did any of the anons download the recent krea2 code lyoko loras for yumi and welita? please share if you have them.
https://civarchive.com/models/2819269?modelVersionId=3183135
https://civarchive.com/models/2819269?modelVersionId=3180138
>>
>>109457725
did i stutter? extrapolate what i said
>>
>>109457705
shes saying gibberish
>>
Hail China for this gift

0 to 3s: the dog on the left which is wagging its tail says "get him, boys". 4 to 10s: as two heavily armed Israeli military soldiers with "IDF" on their chest armor break through the wall on the right side of the room, and grab the man with glasses, taking him out of the room.

https://files.catbox.moe/9n39qt.mp4
>>
File: 8kl8sy.gif (935 KB, 201x202)
935 KB GIF
generate some will smith videos for the next thread
>>
So what date should we hold the official /ldg/ funeral for LTX and Wan? Do all of you guys have your standby obituaries ready?
>>
did I add spectrum correctly?
>>
>>109457776
same day catjack kills herself. we would save so much on the funeral fees
>>
>>109457776
I am writing a heartfelt eulogy for Wan 2.2. Made me coom gallons.
For LTX though, I am only readying my bladder.
>>
Can someone explain to me how IDs work? The Minimax documentation is confusing with this example:

The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The two children (S1,S2) shout together, <d>[English] Wait for us!</d>


Like, who is S1? the "young woman" or the first of the two children?
>>
>>109457785
I have no idea if it combines well with fp16 acc but otherwise seems so.
>>
>>109457785
>>109457801
Also I know you didn't ask these but comfy native int8 loading runs faster than the extension I believe, and heretic TEs are a meme diffusion.
>>
BFL must feel like shit right now. imagine releasing a cucked model right after H3. if they had any sense they'd just cancel flux 3 and spare themselves the humiliation
>>
https://huggingface.co/nvidia/NVIDIA-NemotronLabs-VoiceChat-11B

NVIDIA NemotronLabs VoiceChat is a 11B end-to-end, real-time speech full duplex (FD) model for conversational AI that jointly performs streaming speech understanding and speech generation [1, 2]. Unlike traditional cascaded stacks (ASR LLM TTS), this model achieves full duplex, real-time, seamless voice interaction in one unified architecture, eliminating the need for multiple models or API handoffs, thus reducing end-to-end latency. It sets new benchmarks by bringing open, robust, and highly natural conversation capabilities. Moreover, NVIDIA NemotronLabs VoiceChat is the first open full-duplex model to support tool calling while maintaining a natural conversation flow during tool execution. For each tool, a specific “on-hold” message can be defined that will be spoken by the agent as soon as the LLM generates the text that will trigger the tool call and response.

The model operates on audio signals, which are encoded using a fast conformer module. The resulting audio tokens are inputted into a Nemotron Nano V2 9B LLM backbone to predict text tokens, which are fed to a TTS decoder [2] to predict audio codes for generating the agent's speech. A separate output channel is used to predict tool calling scripts. NemotronLabs VoiceChat offers an unprecedented trade-off between intelligence and latency in the space of open-source voice agents, as highlighted by our benchmarking results below.
>>
So what optimization do you guys prefer, Easycache or Spectrum? What do you use and why?
>>
>>109457789
i think you are more likely than "catjack" to leave this world considering you spent your christmas eve, day, and holiday spamming 4chan; your lover left you alone at the top of mt fuji to go down with another guy; you've been unemployed for over two years and unsuccessfully begging for money everywhere (including your bull whom you aren't even allowed around!); and you spent 2 years spending your energy samefagging without picking up a single user
i'm sorry for you bro
>>
File: 669.gif (3.86 MB, 400x532)
3.86 MB GIF
>i think you are more likely than "catjack" to leave this world considering you spent your christmas eve, day, and holiday spamming 4chan; your lover left you alone at the top of mt fuji to go down with another guy; you've been unemployed for over two years and unsuccessfully begging for money everywhere (including your bull whom you aren't even allowed around!); and you spent 2 years spending your energy samefagging without picking up a single user
>i'm sorry for you bro
>>
holy meltdown lol
>>
>>109457844
EasyCache rapes it as always. Spectrum feels like magic though.
>>
Should I dl fp8 scaled or int8 convrot?
>>
Does the 19gb INT8 pruned model is enough for everything ? Got a feeling its too small to contain all the datas
>>
>>109457863
int8
>>
>>109457867
gm saar
>>
https://files.catbox.moe/inhcur.mp4
>>
>>109457867
the rest of the data is stored in your balls
>>
Honestly, Minimax is blowing my mind. After being cucked to 5 seconds on wan for so long, it felt like we'd be waiting forever to get natively longer videos. Then we got LTX, but it was audiovisual diarrhea and the prompt adherence was always terrible.
Then suddenly minmax h3 comes along, supporting 15 seconds, outstanding prompt adherence, can run on consumer hardware (16gb vram/64gb dram), very little censorship out of the box and already faster than wan2.2 with lightx2v.
The July-August period truly is like christmas time for video generation.
>>
>>109457867
>Does the 19gb INT8 pruned model is enough for everything
Yes
>Got a feeling its too small
Good news is that AI models don't work on arbitrary vibes
>datas
Data is already plural ESL-kun
>>
>>109457877
>supporting 15 seconds
doesnt it support up to 30s 2k? it just takes a lot of time
>>
>>109457877
LTX is good if you know what youre doing.
t.expert at LTX
>>
>>109457856
>EasyCache rapes it as always
Elaborate?
In gen time or quality of output? I was under the impression Easycache takes a heavier toll on quality.
>>
>>109457877
sorry but I don't believe opinions by nogens
>>
>>109457887
based expert
>>
>>109457885
It theoretically supports infinite length, there's not been any reports of it breaking after any length that I know of.
In any case this is the first model that would be able to gen literal movies anyways regardless of the clip length, the reference model is insane and can extend clips out of the box with vid refs.
>>
>>109457844
Of course Spectrum. EasyCache changes the output too much. Also kills the audio. Spectrum doesbn't.
>>
File: detailed video prompt.png (389 KB, 3060x1174)
389 KB PNG
It's impossible to write all this shit yourself, you need a workflow which creates it for you.
I'm working on doing that for the first time, but I'm surprised it isn't more common. You can just load a giant text model to do it, then unload it, you don't need extra VRAM.
>>
>>109457922
you don't need all that shit, basic 1-2 sentence instructions work, look at comfy example
>>
>>109457887
>t.expert at LTX
thats like being the expert in picking the best indian street food vendor, you'll be eating shit regardless.
>>
>>109457922
Why you couldn't write all that shit yourself?
>>
>>109457888
Your impression is correct, I meant in quality.
Easy Cache is faster (with Comfy default parameters) but it very noticeably degrades quality to a point where I find it unusable. Spectrum is moderately slower but quality is a lot higher. Spectrum (extension default settings, the safe profile) makes very minor, almost unnoticeable, changes.
>>
>>109457936
are you kidding me
>>
the possibilities are endless. this is only at 0.3mp too for faster iteration. if you say they fly through a wall, they do.

https://files.catbox.moe/rbg2ek.mp4
>>
>>109457931
I dont make bollywood and hollywoodslop and only make porn with it. LTX is very capable for doing so.
>>
>>109457940
Maybe almost unnoticeable is an a bit overkill description but the delta is impressively low.
>>
File: 46476.png (519 KB, 949x1564)
519 KB PNG
>>109457922
the kingdom of kino is only available to those who work for it
>>
I'm glad this shit still manages to run on 8gb vram+32gb ram and produce good results but I'm really itching for an upgrade now
https://files.catbox.moe/i034ey.mp4
>>
>OOM when trying 15 secs at 0.8MP

Its Over. Not even 64gb of ram + 5070ti is enough
>>
File: Krea2_turbo_00205_.png (3.6 MB, 1448x2176)
3.6 MB PNG
>>109457922
Regular descriptions work but a local ai model should be fine as well
>>
File: 1782981965686281.jpg (95 KB, 1280x720)
95 KB JPG
I am the only person in this world who amidst H3 release have queued 10 days worth of Chroma v48 gens.

What unique /ldg/-related records does anon probaly hold?
>>
File: Krea2_turbo_00207_.png (3.7 MB, 1448x2176)
3.7 MB PNG
>>109457960
without kroma lora
>>
Does order matter for various optimizations if you're adding Spectrum? I'm doing
> sage attention patch
> minimax mem eff attention
> minimax sigma shift
> spectrum
>>
>>109457964
>What unique /ldg/-related records does anon probaly hold?
>18821 Illustrious gens
>36679 Anima gens
>>
H3 can only do realistic, 3D CG, and animu styles. Change my view
>>
it gets more and more impressive when you realize all the stuff you can do.

https://files.catbox.moe/tg8k46.mp4
>>
>>109458003
Yes, its hard to combine all of those, like smooth animation on 2d animu styles. We need loras for this
>>
could I ask for someone to take screenshot of their task manager when they are generating a video on 3090? For some reason, my gens are really slow...
>>
>>109457964
~200000 ZIT gens with 10% of them kept
>>
File: 4754298.webm (3.1 MB, 448x256)
3.1 MB
3.1 MB WEBM
>>
>>109457940
same experience here also, spectrum is tjhe best
easy cache is total trash
>>
>>109457960
>>109457983
kill yourself
>>
>>109457964
Retrained the same lora 30 times with different settings and the best version was... the first one.
>>
>>109458034
uh oh
you sound buttmad
>>
File: Krea2_turbo_00210_.jpg (3.13 MB, 1448x2176)
3.13 MB JPG
>>109458034
+1
>>
>>109457936
Sir my English is not so good, why would I type many words by myself if I can ask ai to do the needful and write beautiful poem from my rough explanation?
>>
>>109458017
0.4mp vs 0.3 (prior)

https://files.catbox.moe/nykvhx.mp4
>>
Damn spectrum is good, shaves like 30-40% off of gen times for little loss in quality or motion.
>>
>>109457964
Manually tagged with simple tags 40000 images I generated by now.
>>
>>109458071
Youre lying
>>
https://github.com/Haoming02/sd-webui-forge-classic/issues/1380
>>
>>109458067

Yes, Sir,. the prompting is Parfect for gorgous Looks in the AI model.
>>
>>109457995
>sage attention patch
>minimax mem eff attention
These two are exclusive, if both are applied, only one should take effect, remove the regular sage.
I am testing how it combines with attn optimizations, can't report on how well it works for now.
>minimax sigma shift
This is not an opt, just parameter tuning
Possibly worth noting it won't combine with easycache.
>>
https://files.catbox.moe/cp6bp8.mp4
Nipple lora when? I can't take it anymore
>>
>>109458089
>This is not an opt, just parameter tuning
But original comfy workflow doesn't have this node
>>
>>109458092
Just commit suicide already. You've been here for ages and nobody likes you, idiot.
>>
File: MiniMax_H3n_00063.mp4 (1.45 MB, 1376x768)
1.45 MB
1.45 MB MP4
>>
lodestones just announced that he's working on a h3 finetune, cant wait
>>
>>109458092
i enjoyed the demonic piano playing ion the background
>>
Where did it all go wrong for Framepack?
>>
File: 1768895435536793.png (452 KB, 736x416)
452 KB PNG
based China made this possible.
>>
>>109458104
Idk, your mom likes me
>>
Pajeets grab Lara Croft's bobs and vegana
https://files.catbox.moe/9ct3k8.mp4

H3 is still sloppy. any way to improve quality?
this was at 0.6mp used rtx upscaler too
>>
File: 1773243075279951.png (861 KB, 2107x1153)
861 KB PNG
https://i.4cdn.org/wsg/1785851301114689.mp4
>>
>>109458124
bodied that freak
>>
>>109458101
Yes, it uses the default shift values (12 for video and 3 for audio) if you omit it. The node allows you to overwrite them. It has no effect on the speed of generation, just output quality. Read comfy_extras/nodes_minimax_h3.py
>>
>>109458154
MR ADVERTISER GET DOWN!!!
>>
>>109458154
I have hereby decided not to advertise my product on 4chan
>>
File: MiniMax_H3_00005_.webm (2.54 MB, 544x800)
2.54 MB
2.54 MB WEBM
lol I got permabanned from reddit for posting this video (by reddit admins)
>>
>>109458154
free 3 day vacation
>>
File: 5466262.jpg (294 KB, 1026x595)
294 KB JPG
is this normal for h3? generating 10sec 480p video
>>
haha wtf it made the video like a tiktok/youtube clip or something without me prompting the sounds.

https://files.catbox.moe/me1ysa.mp4
>>
>>109458154
are you prompting the boob physics or is it doing that all by itself?
>>
File: download.jpg (46 KB, 600x603)
46 KB JPG
are we behaving ourselves, anons?
>>
>>109458169
>64gb
ramlets get out REEEEEEEEEEEEEEEEEEE
>>
HAHAHA

holy shit, same prompt but diff seed. id watch this show.

https://files.catbox.moe/cmeksa.mp4
>>
File: Krea2_turbo_00225_.png (3.8 MB, 1400x2096)
3.8 MB PNG
>>
Minimax can't generate nice looking panties on anime characters. They need to be part of the ref.
>>
>>109458232
It can
>>
>>109458231
Aren't you the retard that had an existential crisis because some anon baked with will smith eating spaghetti gif was the OP?
>>
Now that the dust has settled, I think LTX won.
>>
>>109458240
Prove it. Been prompting for a bit and the panties look kind of shit and lack cute/sex appeal. Kind of like male underwear.
>>
File: 1782182198963800.png (76 KB, 300x265)
76 KB PNG
Should I be applying the "MiniMax H3 Mem Eff Sage Attention Patch" and "Spectrum Apply MiniMax H3" node to model and inputting into both BasicScheduler and Basic Guider? Or just Basic Guider and have the Scheduler without them? Instructions unclear. Google says do both.
>>
File: Krea2_turbo_00229_.png (3.78 MB, 1400x2096)
3.78 MB PNG
>>109458231
Higher steps give more details lower seems more uniform.
I need to test this more
>>109458249
Have you seen any of my gens?
If not then I wasn't on. I don't spend all my time here. Will you ever get over the fact everyone hates your camp?
>>
>>109458255
genning it right now. Give me 2 minutes
>>
>>109458249
that avatarfag in particular has been melting down for years and is the schizo baker that made the schizo rentries
>>
Fuck, the shit-grade audio I had in videos is literally frpom beta scheduler. Never change it - leave it on simple guys...
>>
>>109458251
...the losing competition.
>>
>>109458089
>>109458140

Noted, thanks anon. Do you (or anyone else) have feelings about good sigma shift settings, I'm just on default
>>
https://files.catbox.moe/i19lgh.mp4
she's having a lot of trouble finding a comfortable spot in her chair
>>
>>109458259
>Have you seen any of my gens?
>If not then I wasn't on. I don't spend all my time here. Will you ever get over the fact everyone hates your camp?
Can someone translate this schizo? I have no idea what this has to do with his meltdown
>>
>>109458259
>1girl, big boobs,portrait since 2023 straight
your gens are meaningless and ugly. you are not even protesting at this point.
>>
File: ws.jpg (161 KB, 1796x1026)
161 KB JPG
how is he so powerful?
>>
>>109458258
Technically it doesn't matter for the scheduler.
You should do both though as it has no harm and keeps your flow cleaner.
>>
File: H3_00003.png (578 KB, 864x480)
578 KB PNG
finally, a model that makes me regret not buying a 5090
I think. People's faces from far away are still kind of gross but maybe that's a resolution/quantisation issue
>>
the good thing about h3 is that it only needs a general sex/nudity lora, because it already knows so much. unlike wan 2.2, it doesn't need a separate lora for every sex act and facial expression
>>
>>109458249
>some anon
That was obviously you faggot

Whenever local gets a new powerful model there's someconcerted attempt at derailing /ldg/ threads

Go back to /sdg/ and seethe
>>
>>109458268
4 for video is good. i don't know why they chose 12
>>
>>109458287
your level of seethe tells me otherwise
>>
How is minimax with anime character knowledge? Does it know any anime characters by their appearance and/or voice? And if so, which are the notable examples?

What about ref2va? Anyone tried it yet using a voice ref? Is it similar to using a voice ref for a tts model like qwen3-tts and voxcpm?
>>
File: chernobyl 1.png (1.37 MB, 1024x1024)
1.37 MB PNG
>>109458003
>Have to go back to LTX with loras because minimax can't do pixar characters with realistic human like movement, it always moves like a bad cartoon no matter how much proompt correction I do.
I got scammed
>>
File: 1770914823753041.webm (2.19 MB, 960x1440)
2.19 MB
2.19 MB WEBM
>>109458255
Here you go
>>
>>109458268
I am focusing on testing attention optimizations first, no shift experiments from me yet, sorry.
You probably want to just leave it on default for now.
>>
>>109458003
>H3 can only do realistic, 3D CG, and animu styles. Change my view
h3 can only do styles people care about, yes. if you want something niche, use a lora
>>
>>109458298
NTA but those are shorts
>>
>>109458278
He's the butt of a long-running joke

To top it off his wife publically declared him a cuckold
>>
>>109458294
just from text2vid prompting its anime knowledge is very lacking. but with references anything is possible I suppose
>>
File: MiniMax_H3n_00067.mp4 (1.82 MB, 1376x768)
1.82 MB
1.82 MB MP4
>>
>>109458297
Go away with your jewish tricks
>>
>>109458298
You just proved my point. Look at those hard black outlines, and that simple crease going down the middle. Those are not good panties, they are video-gen panties. If this were image gen they would definitely need to be inpainted.
>>
>>109458305
Give me 20 minutes
>>
File: 08901643.mp4 (3.5 MB, 448x672)
3.5 MB
3.5 MB MP4
>>
File: file.png (187 KB, 1042x789)
187 KB PNG
>>109458293
nta but you will never be credible saying anyone is "seething" when you have publicly spent your entire christmas having a melty here
i'm not sure how you aren't mortified showing your face here lol
>>
Will Minimax be lora friendly? There are still 0 loras on civitai
>>
>>109458310
Do I remember the lore correctly, she was some chick from a wealthy family who did porn ?
>>
>>109458321
>>
>>109458327
you are having a meltdown right now faggot. go put the fries in the bag lilbro
>>
>>109458336
that really doesn't hurt coming from a 3 years unemployed and e-beggar reject lol
>>
>>109458330
It needs an adapter to even be trained properly.
Some very early support by ostris and that's it for now, we'll probably start to see proper loras in like a week or so.
>>
ostris is working on a h3 training adapter
https://xcancel.com/ostrisai/status/2084642732610396411
>>
>>109458330
Civitai pretty much always have a 5-7 day delay before loras start trickling in for a new model

That said wasn't Minimax lora training impossible until yesterday, when AI Toolkit added support ? Have patience young padawan
>>
>>109458346
someone made an experimental musubi fork for training it.
>>
File: MiniMax_H3n_00064.mp4 (2.03 MB, 1376x768)
2.03 MB
2.03 MB MP4
>>109458332
yes
>>
>>109457922
people really out here boomer prompting like it's the good old 1.5 days again instead of realizing you can just write 2 succinct sentences and get the same result
>>
>>109458316
I'm not saying its bad, it's just disappointing when people here are glazing the model like it's the second coming of Christ. I'll just wait for H3 loras.
>>
>>109458370
>you can just write 2 succinct sentences and get the same result
no wonder ai gens are soulless now
>>
I would like to formally apologize for mocking and disrespecting Chinese Culture.
>>
File: 1754909285916124.webm (2.99 MB, 960x1440)
2.99 MB
2.99 MB WEBM
>>109458318
>>109458305
Only took 500 seconds. I need to install those Spectrum everyone talking about
>>
using a video in ref flow seems to not do anything to change the video, is it a resolution thing? i'm following the prompting guide, at least i think.
>>
I'm a tard when it comes to generating images, how easy is the chink h3 model to use? Really have no clue how you guys use comfy so well, shits messy.
>>
File: Krea2_turbo_00261_.png (3.32 MB, 1320x1984)
3.32 MB PNG
>>109458287
Kind of hilarious seeing him still here, he never had the makings of a functional member of society
>>
>>109458403
still has hard seams. these are similar to boys underwear.
anon go and google female underwear refs and anime pantyshots then start prompting again.
>>
>>109458403
>Only took 500 seconds
because you're using 1 megapixel
>>
Is DramaBox TTS still good? I want to gen audio for the audio ref.
>>
>>109458461
Whats what panties are look like anon.....

>>109458467
0.6mp with RTX super reso 1.5x. I got OOM when trying 0.8+mp at 13+ seconds
>>
File: MiniMax_H3_00107_R.mp4 (3.34 MB, 768x1376)
3.34 MB
3.34 MB MP4
>the main subject is the very tall woman from <picture 0>, she is 3 metres tall,
>>
>>109458453
you bought rtx 5090 for spamming this?
>>
>>109458471
yes
>>
>>109458423
Just use the default workflow.
Technically you are supposed to put some effort into prompting but it responds reasonably well to lazy tard prompts.
>>
>>109458471
I was experimenting heavily with tts before minimax dropped. The ones you want to try are Index-tts2, omnivoice, qwen3-tts, Fish Audio S2 Pro (very slow)
>>
>>109458478
>0.6mp with RTX super reso 1.5x.
this is a smart idea anon, can I see your ComfyUI nodes that you used to accomplish this if you're doing this in comfy
>>
File: Krea2_turbo_00263_.png (3.3 MB, 1320x1984)
3.3 MB PNG
>>109458481
I plan to buy a bigger gpu once nvidia stops being cheap with vram with the rtx pro.
Having a job allows you to do that
>>
>>109458508
Just put the nodes after VAE decode
>>
https://files.catbox.moe/rvhtbs.mp4

kino on loop
>>
Anyone messed around with audio references yet? I've had decent results using a 15 second clip as a timbre reference, it's not perfect but it's a huge improvement over the model's default output
>>
>>109458471
>>109458482
https://huggingface.co/ResembleAI/Dramabox

Do you guys have functioning ears? Can you not hear the glaring issues with the examples on this page? Sounds like dialogue coming through a tube or a toilet roll.
>>
>>109458524
shut up goy
>>
h3 makes some kino music
>>
>>109458523
I kept getting garbled audio at the end of a voice line with long clips, I don't think they've offered guidance about the audio, so I was just using the clips I had for VibeVoice.
>>
How viable is minimax for actual storytelling? Is there some kind of pipeline that can maintain consistency between environment, characters, clothing, voices and composition?
>>
>>109458512
Money can't buy you any sense or creativity though.
>>
>>109458565
You can probably ask visual LLM to prompt every next 15s part based on the current last frame and your overall plot, but this was surely done long ago with older models
>>
>>109458565
It's too new, soon there will be a gazillion youtube videos and workflows for this very purpose
>>
>>109458565
At the very least I'd expect being able to use three different reference videos would allow you to highlight elements that you want preserved from one gen to the next. Out of the box that should be a huge improvement on ltx or wan.
>>
>>109458569
Zero self awareness does that to a lolcow
>>
File: post.png (217 KB, 586x539)
217 KB PNG
>>
https://files.catbox.moe/bvrjhc.mp4
>>
Hey, I'm new to this shit and I would like to just create a video where I replace my own face and voice in a plain background video with my friend's face and voice. I have long videos of my friend talking directly to a camera for training material, so that is not a concern. Could anyone point me in the right direction of how I could do that? Thanks.
>>
>>109458565
I think you can this model is so much better, I just wish I had more drive because I have two finished scripts and I don't know what to do with them. It is an insane amount of work for a person but an llm can do a lot of heavy lifting with the prompts and such.
I think the key is using references for everything, characters, backgrounds, etc. What would require more expirementation would be prolonging a scene out of different gens.
>>
>>109458614
>Slow motion turbo loras
No, I won't do it again
>>
File: waiting.gif (450 KB, 128x128)
450 KB GIF
>>109458614
KINO!!!!!!!!!!!!!!!!!
>>
Good morning, using Miniman H3, how do I prevent the SamplerCustomAdvanced node from failing?
And ComfyUI fell hard, holy shit. It was a different GUI 1 year ago, now it's a corporate dashboard. I hate it.
>>
>>109458412
ok looks like using <video 0> and <subject 0> are wrong, if you start at index 1 things start working.
>>
>>109458639
How do you expect anyone to provide help without telling what error you are getting?
>>
>>109457799
Anyone?
>>
File: Krea2_turbo_00278_.png (3.01 MB, 1320x1984)
3.01 MB PNG
>>109458614
Big things are happening
>>
File: 1759196720922354.webm (1.54 MB, 960x1440)
1.54 MB
1.54 MB WEBM
Im gonna smooooooooke



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.