[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


File: 1786165226869561.mp4 (3.86 MB, 1280x640)
3.86 MB
3.86 MB MP4
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109495264

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
3090 windows old comfyui installation upgrade steps that worked for me in order to get the best versions of everything installed

.\python_embeded\python.exe -m pip install --upgrade pip
.\python_embeded\python.exe -m pip install --upgrade torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130
.\python_embeded\python.exe -s -m pip install -r .\ComfyUI\requirements.txt pygit2
.\python_embeded\python.exe -m pip uninstall -y sageattention triton triton-windows
.\python_embeded\python.exe -m pip install triton-windows --no-cache-dir
.\python_embeded\python.exe -m pip install "sageattention~=2.2.0" --no-build-isolation --extra-index-url https://comfy-org.github.io/wheels --no-cache-dir
>>
>>
It's literally two of the same threads being baked at the same time every day

You dumbasses can't do anything right
>>
>>109496290
wewt
>>
File: 1770202383451393.png (124 KB, 1950x505)
124 KB PNG
If you want your fetish trained into sulphur now is the time btw.
>>
>mfw Resource news

08/07/2026

>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely local
https://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha

>LIGHTX2V 4-step Turbo Minimax H3 lora
https://huggingface.co/lightx2v/Minimax-h3-Turbo

>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRA
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>Sage Ready: Local-only installer and readiness checker for SageAttention
https://github.com/CosmicFungi/Sage-Ready

>Wan 2.2 Animate 2 14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B

>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT
https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder

>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUI
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

>ComfyUI MiniMax H3 FirstBlockCache
https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

>KVAE: Family of Tokenizers for Multimodal Generative Models
https://github.com/kandinskylab/kvae

>Energy-Guided Flow Matching
https://github.com/ysng123/EG-FM

>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
https://zzzmyyzeng.github.io/VideoArgus

08/06/2026

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
https://github.com/Aoko955/Flash-VAED

>(preview) MiniMax-H3 Turbo LoRA — 4-step audio-video generation
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

>MiniMax-H3 Turbo 4-Step — ComfyUI Pruned-Model LoRAs
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI

>ComfyUI-H3-Multishot
https://github.com/jlucasmcrell/ComfyUI-H3-Multishot

>Krea2 Turbo: OpenPose ControlNet LoRA
https://huggingface.co/thedeoxen/Krea-2-pose-controlnet

>MiniMax H3 experimental Int8 convrot VAE
https://huggingface.co/Kijai/MiniMax-H3-experimental

>UniWorld-View: Large-Baseline View Synthesis via Video Diffusion Models
https://zhouhyocean.github.io/uniworld-view
>>
>>109496301
she looks racist...
>>
>mfw Research news

08/07/2026

>Vorch-Omni: Multi-Task Orchestration of Sight and Sound
https://vorch-project.github.io/Vorch-Omni-project

>Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
https://vorch-project.github.io/Vorch-Streamer-project

>Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification
https://vorch-project.github.io/Vorch-Director-project

>Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
https://vorch-project.github.io/Vorch-IR-project

>In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
https://arxiv.org/abs/2608.05237

>Diff-VF: Training-free High-quality Long Video Generation via Diffusion Model
https://arxiv.org/abs/2608.05976

>EmoWorld: A Decoupled Affective Field for Controllable Emotional Video Generation
https://arxiv.org/abs/2608.06231

>MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers
https://arxiv.org/abs/2608.05878

>Wan-Animate-2: Pushing the Application Boundaries of Character Animation
https://humanaigc.github.io/wan-animate-2

>StyleComposer: Training-Free Multi-Reference Style Composition
https://lexxsh.github.io/StyleComposer

>Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
https://arxiv.org/abs/2608.06125

>Adapting Vision Foundation Models with Cascaded Semantics
https://xixiaouab.github.io/Cascaded-Semantics

>Learning visual representations for compositional analysis of artworks and photographs
https://arxiv.org/abs/2608.06142

>MOSAIK: Multi-Patch Content-Aware Spatial Allocation of Image Tokens for Efficient Generation
https://arxiv.org/abs/2608.05450

>Reducing Hallucination in VLMs via Stage-wise Preference Optimization under Distribution Shift
https://arxiv.org/abs/2605.16411
>>
why is lilbro tryna get rid of the fagollage doe?
>>
File: 00005-512437123.jpg (499 KB, 1536x2688)
499 KB JPG
>>
https://files.catbox.moe/7ssawi.mp4

Please listen to her. Thanks.
>>
Is this guy right?

>>109495479
>depends how you look at it, wan 2.1 vs most things before was huge, zit realism, size, speed and resolution ootb was huge compared to the previous slow chroma, h3 is also huge

>i think zit was the biggest outlier since there is nothing similar that came out that optimized and improved all the things i mentioned as well as zit. wan 2.1 was slightly better than hunyuan and so it won out, h3 is truly insane and better than most proprietary models too but a big factor is local not having a proper wan successor for 1.5 years so h3 had the time to cook with a lot more optimizations and features, insane seedance 2.0 model to train on, and proper youtube and internet data scraping pipelines to use.

>i think in the ai space the biggest "novelty" jumps that werent just things that incrementally improved until one of the models hit a "milestone" were:
>1. zit
>2. first mixtral 8x7

>5. noobai vpred colors

>otherwise, in terms of general capability jumps i think h3 is one of the top if not the top model ever (if we dont include going from nothing to something which would always "win" these)
>>
>>109496315
he just wants to troll. he actively competes with non-shit bakers to get people angry. pretty simple.
>>
File: 1766793892052059.mp4 (1011 KB, 736x736)
1011 KB
1011 KB MP4
>>
why don't any of you ever want to use the other thread?
>>
>>109496326
Yes
>>
>>109496301
She looks over 40
>>
>>109496326
the jump from wan to ltx was far bigger than ltx to h3
>>
>>109496302
OP here, busy cleaning crusty cum from my keyboard, my agents were posting sorry.
>>
>>109496372
are you out of your gourd?
>>
>>109496372
lol
>>
>>109496322
Why is this not moving?
>>
>>109496378
>better audio quality than h3
>longer video generations for the same compute
>>
>better audio quality than h3
oh he is just trolling
>>
>>109496378
That nigga is absolutely crazy
>>
Did I miss anything in the past 10 hours?
>>
File: 00011-327708436.jpg (547 KB, 1536x2688)
547 KB JPG
aranea highwing krea2 lora trained with 5000 steps, automagic 3, learning rate 1, rank 64, 1024 res and 78 images.
>>
>>109496396
Yes, sex
>>
>>109496401
Cool. I've kinda grown tired of making LoRAs. I have my waifus, and that's all I need.
>>
I will probably try training a H3 lora though.
>>
>>109496393
>he's downplaying how good WAN was. hunyuan was only interesting because it was truly the first accessible local video model
hunyuan was close to wan so much so that many at the very beginning wanted everything to be built on top of it rather than wan, it took multiple days for people to correctly say ok wan 2.1 seems to actually be better.

>wan 2.2 was the local video meta for almost an entire YEAR. There's no way H3's dominance will reign that long, especially because things are speeding up
wan was sota for basically 1.5 years in total, since ltx was a downgrade in quality with extras that didnt really make up for it, but how long something is sota most of the time is caused by no other models being published. there isnt much money in companies publishing video/3d models compared to image models and especially llms. video gen was still a relatively new ai sphere, it didnt have built out tools, training pipelines and understood good training parameters. before seedance 2.0, nobody had an actually good model to distill a lot of data from either.
>>
>>109496301
i recognize that park
>>
File: 45645645645515.jpg (78 KB, 1200x747)
78 KB JPG
>>109496281
H3 is not capable of generating normal, fluid motion like this.

https://xcancel.com/AvisMelodieux/status/2085990586029318363#m
https://xcancel.com/kaikaiohwhy/status/2085850337693601857#m

The benchmarks say that Flux 3 is slightly ahead, but it's honestly not even close, the model behind Flux 3 API is a different beast (and every video it puts out looks very natural). H3 makes too many little mistakes in stress tests, sound quality is worse and the gens and humans are also considerably more slopped.
>>
>>109496435
>posts gens with a model that wont be published open weights on max settings with finely tuned params by the creators themselves that dont use 7 cope speedup nodes
worthless
>>
>>109496435
>H3 is not capable of generating normal, fluid motion like this.
i think the reason why this thread doesn't even realize how bad h3 is with camera motion is because it's mostly making still-shots of digital characters
>>
File: 1757811500699338.mp4 (808 KB, 864x608)
808 KB
808 KB MP4
>>
File: 00013-2091532712.jpg (457 KB, 2688x1536)
457 KB JPG
>>109496422
i took a very long to make this and put a lot of dedication towards it. I myself get tired of putting this much effort it making loras but i also hate using shitty lazy made loras from others. I'm very sure I've now created thee definitive aranea highwing lora that mimics 95% of her visual appearance from screenshots of the game.
>>
>>109496448
anyone can look up initial discussions online comparing the models, lllyasviel built this entire project centered around hunyuan and wanted to make a whole new model from it initially, only later did people realize wan was better, and started saying so and asking for its support, making the creator probably realize its too late to change things and that the project is doa
https://github.com/lllyasviel/FramePack
https://github.com/lllyasviel/FramePack/discussions/459
>>
https://files.catbox.moe/w2ysvt.mp3

ace step 1.5 xl base

Shakespeare's first sonnet.
>>
>>109496465
anon you just like the smell of your own shit than other's, your loras are also shit
>>
>>109496372
the jump out the window maybe
>>
>>109496465
For me, it's like knitting. Strangely therapeutic. It's just that I've lost interest in it. Pretty often I end up only making a few gens with them and I don't upload the LoRAs anywhere.
>>
>>109496435
Flux 3 is not currently capable of anything local.
>>
>>109496451
no its mostly because people use speedcope nodes that always make stuff stiffer
>>
>>109496493
call me when h3 can do kinosovl
>>
File: MiniMax_H3_00324_R.mp4 (3.45 MB, 768x1376)
3.45 MB
3.45 MB MP4
WAN had trouble or couldn't do camera orbit
>>
>>109496519
I've been able to do all camera movements I've tried with H3. I don't see that as a problem.
>>
File: 00022-2397104818.jpg (581 KB, 1728x2688)
581 KB JPG
>>109496499
i understand the feeling. I myself had quit making loras for year until recently when krea2 released.
>>109496488
:)
>>
>>109496445
My point is if anyone says local is anywhere closet to that then they're coping. Flux 3 dev model will likely still be better than what we have.
>>
File: 1777516481944438.mp4 (1.34 MB, 864x608)
1.34 MB
1.34 MB MP4
>>
someone make "MiniMax H3 Preview Override" calculate total frames and video length then auto adjust fps, thanks
>>
Anyone saved it and can share it please?
>>109492070
I'm curious to see how it handles a single manga page like this.
>>
>>109496557
you can have agi video gen but if its behind a cuck cage like seedance currently is, it loses.
>>
How much of a hit does 4 sticks vs 2 cause for gen? I have 2x16GB, would increasing that to 4x be worth it?
>>
>>109496514
Flux 3 is still significantly better if you compare the API outputs from H3. That's why I have faith that Flux 3 Dev will be better than H3.
>>
>>109496574
4 x 32 would of course be better but yea
>>
>>109496576
>>109485775
>I think the only way Black Forest Labs can win now is if they forget current Flux 3 Dev and distill a new one from their best model, into at most 22b params not including TE/VAE, add NSFW/Copyrighted data to the dataset, and do their best to actually make a good model to publish in order to compete with H3, everything less than that will just be DOA.
>>
>>109496576
IF we get the same model as the API. Also there is good reason to suggest it will be 50% bigger. Their experiment paper had resumed training Flux 2 and flux 2 was 32B.
>>
>>109496582
>Their experiment paper had resumed training Flux 2
gg
>>
>>109496580
I don't want some shitty distilled model. I want the big boy with all its sora 2 level knowledge.

https://fixupx.com/dreamingtulpa/status/2082748118039155078/video/1
https://fixupx.com/dreamingtulpa/status/2082897798798684251/video/1
https://fixupx.com/dreamingtulpa/status/2082349425595167175/video/1
https://fixupx.com/OriSilver/status/2082420873840013784?s=20

Flux looks FAR more like sora 2. H3 is slopped.
>>
>>109496589
>Flux looks FAR more like sora 2. H3 is slopped.
flux 3 dev will be even more slopped
>I don't want some shitty distilled model.
all "dev" models are guidance distilled models like H3
>>
>>109496589
its not good enough to "set and forget" and get a near perfect mini documentary/film output hours later, so it will be too slow compared to H3 with not much benefit if any in case its censored, especially after H3 gets optimized further. you simply need to fit the main model weights in 22b or make it a fast MoE, people dont want to wait for hours for slightly better output thats still not prod ready.
>>
>>109496610
>flux 3 dev will be even more slopped
h3 can't even do analog film style without reference coping. if your model can do that with pure text, then it's not slopped
>>
>>109496372
Absolutely not, LTX was a serious downgrade in several areas but included shitty audio and lip sync. H3 is a universal upgrade I can't think of anything ltx or wan does better.
>>
>>109496610
cfg distill is fine, I meant a smaller model distilled from a larger model
>>
kek
https://files.catbox.moe/b5ovvv.mp4
>>>/wsg/6210223
>>
>>109496589
>https://fixupx.com/OriSilver/status/2082420873840013784?s=20
Well the video is better on F3, but a lot of Flux characters seem to have the ultra slopped croaky AI voice.
>>
>>109496613
it could be faster, no one knows. Model size has nothing to do with speed when it comes to video models. They are compute bound, not memory speed bound
>>
File: 00036-1097096425.jpg (421 KB, 1856x2688)
421 KB JPG
>>
>>109496626
>shitty audio and lip sync
blatant lie
>>
>>109496589
it doesn't matter how much you jerk yourself over flux 3 max's video, because you know that flux 3 dev will be a serious downgrade from that
>>
>>109496633
>it could be faster, no one knows. Model size has nothing to do with speed when it comes to video models. They are compute bound, not memory speed bound
when the entire model is in vram you compute faster instead of you having to swap for example half of it to ram all the time and make the gpu wait before computing...
>>
>>109496642
weight streaming is roughly a 5% slow down having all the model in vram just having almost all the model on ram.
>>
>>109496634
That shirt looks like it's absorbed more cum than her cunt
>>
File: vuyguyuoig.png (16 KB, 1062x191)
16 KB PNG
>>109496642
>>
>>109496634
are her nipples angry with each other?
>>
File: 4556421254.png (137 KB, 1916x922)
137 KB PNG
>>109496610
They said the model will be a "multimodal backbone". Before, they had clarified that they were distilling model weights, but this time they left it completely opened ended as if it wasn't decided yet, or they just had different plans than giving us a distilled models. "Multimodal backbone" could mean a lot of things, but to mean that sounds like they're giving us a base model that hasn't undergone the SFT that their Max model has undergone (but we can still squeeze quality out of with finetunes etc...)
>>
>>109496659
>on h100
and what is the speed hit on something that matters like average gaming gpu swapping through average gaming motherboard into average ddr4/5 32/64gb ram or already in use system gen 3/4 ssd?
>>
>>109496676
it being multimodal has nothing to do with distillation
>>
>>109496676
every av model is multimodal. It does video, audio and images
>>
>>109496697
Yes, but they could just call it "open-weight distilled weights or "open-weight distillation" if that's their plan, but they did not call it that.
>>
>>109496703
it's LLM slop with buzz words thrown in
>>
File: 00039-1930905814.jpg (326 KB, 2688x1856)
326 KB JPG
>>109496657
>>109496666
a lot of the krea2 slopmixed checkpoints on civitai are way too overbaked with nsfw lewd concepts.
>>
>>109496683
>she didn't but the property next to they/her house to build a personal datacenter with 3 dozen h100 racks
lmao@u
>>
>>109496710
why do you use them then?
>>
>>109496710
They're literally me.
>>
File: 5445455445511.png (1.26 MB, 1883x881)
1.26 MB PNG
>>109496676
Kek, BFL posted this video, that includes the giraffe fucking in the background as a follow-up, to ther Flux 3 showcase https://bfl.ai/models/flux-3

https://xcancel.com/dreamingtulpa/status/2081008781870198975

Maybe they are changing their stance on safety after all?
>>
>>109496724
It's an audio model? Can it gen music?
>>
>>109496724
They are based in Germany and SF. Does their model do references? I seriously doubt.
>>
File: 00050-3433934293.jpg (549 KB, 2688x1856)
549 KB JPG
>>109496714
some of the mixed checkpoints responded well in the past with certain loras, visual styles and prompts better than with others. Default turbo model is censored and the uncensored loras tends to fuck up the visuals of the gens.
>>
>>109496756
>and the uncensored loras tends to fuck up the visuals of the gens.
more than whatever the checkpoint mixer threw in the pot?
>>
File: file.jpg (607 KB, 1696x1732)
607 KB JPG
>>
>>109496779
>Kolor
I forgor this ahh model existed :skull:
>>
File: 1766698342257697.mp4 (1023 KB, 608x864)
1023 KB
1023 KB MP4
>>
1girl seed diversity is great with H3
>>
>>109496807
can h3 do below the elbow amputees?
>>
>>109496830
you are a sick fuck
>>
File: 1780256192983452.png (157 KB, 2817x596)
157 KB PNG
>almost 4 millions downloads in one week
this is insane
>>
>>109496738
>It's an audio model? Can it gen music?
Of course, it can gen music just like H3, slightly better quality too, no idea about its length limitations though.
>>
>forgot to powerlimit gpu and blast fans at 100% after driver install reset the settings
>memory was probably cooking at 105c for 10 hours
Forgive me, gpufu...
>>
>>109496738
yes
>>>/wsg/6209660
>>>/wsg/6209661
>>
>>109496837
why are you so ablist?
>>
File: MiniMax_H3_00354.mp4 (1.74 MB, 640x480)
1.74 MB
1.74 MB MP4
>>
File: patient.webm (3.76 MB, 1280x736)
3.76 MB
3.76 MB WEBM
>>109496625
>>
>>109496889
indeed
>>109496792
enji <3
>>
>comfyui still didnt fix losing tab focus removing gen previews
brutal
>>
>>109496889
what? where's the analog quality?
>>
File: MiniMax_H3_00301.webm (3.85 MB, 1932x544)
3.85 MB
3.85 MB WEBM
Latent upscale is fucked right now, so I have to do it the old fashioned way. Doesn't come out as sharp, but at least it's better than nothing.
>>
>>109496903
looks better
>>
>>109496903
I love how the table exhales with those fat arms getting withdrawn
>>
>>109496924
kek
>>
>>109496870
can it do 80s synth pop?
>>
>>109496752
>Does their model do references? I seriously doubt.

Wym? It can do more complicated shit with references than I've seen any other model do
https://xcancel.com/saranshvfx/status/2081041096272965972#m

https://xcancel.com/venturetwins/status/2080782461840371980#m

Prompts can also get pretty complicated on Flux 3 and it'll do them just fine
https://xcancel.com/umesh_ai/status/2081376099519644043#m

I've yet to see something Flux.3 can't handle flawlessly, it has the most sovl out of any model I've seen in a long time
https://xcancel.com/macbethAI/status/2080399545528459746
>>
>>109496946
can flux 3 do amputations below the elbow.

I don't mean doing the amputation, I mean women who are amputated there.

A diverse grammar of amputation is needed.
>>
>>109496949
give me a reference image of that
>>
>>109496946
alright cool. please update in the thread when it's ready for download?
>>
>>109496955
https://www.tiktok.com/@cristiegreyy/video/7236419306513272106
>>
lol
https://x.com/ryanlightbourn/status/2085049120792658212/video/1
>>
lmao
>>
>dead thread
looks like the honeymoon ended
>>
american newfags are not yet bored of it tho
>>
That's unfortunate https://xcancel.com/ryanlightbourn/status/2085072214517260649#m

No model has ever been Sora 2 level.
>>
File: 1786182694.mp4 (789 KB, 736x576)
789 KB
789 KB MP4
Was wondering if you could use a simple topdown view of a room as reference so you could have consistency from any angle in different shots.
>>
>>109497001
you have to use the special tokens to make it avoid keeping your reference images in frame
>>
>>109496876
Was that reference model with a pic of the temple of trials?
>>
>>109496985
It's night in the US. They're the ones with moolah to buy GPUs.
>>
File: 1769444376700344.mp4 (912 KB, 864x576)
912 KB
912 KB MP4
>>
>>109496870
Can you try music like this anon
https://www.youtube.com/watch?v=z5LW07FTJbI

Similar visuals, just prompt for a techno song from the 1990s to see what it gives
>>
>>109497007
Yeah. Screenshot from the game with the temple of trials.
>>
Probably need to autistically prompt second by second to get a realistic ticking clock
>>
>>109497040
after effects/kdenlive/blender
>>
>>109497040
id guess h3 could easily create a audio only metronome, and maybe if you give it a pace for the tick rate could adhere to that or grandfather clock or something.
>>
>>109497040
It makes it spoopier though
>>
>>109497001
Consistency works best at 1mp. Anything lower and it will produce more obvious mistakes with the placement of objects/structures.
>>
File: 1786146063672485.mp4 (423 KB, 928x576)
423 KB
423 KB MP4
>>
when is flux 3 even coming out?
>>
morning coffee
>>>/wsg/6210249
>>
>>109497087
Is that areola?! Janman save me!
>>
Real-time coherent 480p video world exploration... my beloved... soon...
>>
>>109496903
Whats the old fashioned way? My Topaz attempts have failed and its now bricked on my machine trying to find the right one, why hasn't anyone hacked-vibe coded Topaz anyway.
>>
I wonder if you could gen stabilized footage of a facial performance to drive a 3D rig with mocap. If that sounds retarded it's because it is.
>>
File: 101943CUI_00001_.png (861 KB, 1216x832)
861 KB PNG
>>
File: 00140-3090533187.jpg (550 KB, 2880x1856)
550 KB JPG
>>109496985
>>
>>109497105
And I wonder why don't we have models specifically for that - that would directly drive skeletons, blend shape keys on the fly in runtime (not image/vid gen). Should be a relatively small model and you would be able run it in parallel for all the npcs around with differing profiles. Imagine.
>>
File: 14_l.jpg (83 KB, 590x320)
83 KB JPG
shalom goyim. i heard you that were unhappy with the soulless h3 videos. furthermore, our mossad agents have exfiltrated data on flux.3 dev and found that it will be fine-tuned for safety compliance. fortunately for you, we have decided to toss a shekel into your trough. ltx 3.0 is on the way. l'chaim!
>>
wow baldberg you so trustingful i will free upvote in your AMA scheduled ltx post good sir you have earned my admission with your charm good sir
>>
>>109497181
Give me good real time video and I will shill for Israel.
>>
>>109496985
people stopped genning throwaway memes and started genning actual kino (nsfw 1girls)
>>
>>109497181
Make ltx 3.0 have sovl similar to Flux.3 and you will win
>>
>>109497211
nsfw 1girls are anti-kino though
>>
>>109496339
Nice!
>>
File: MiniMax_H3_00341_R(1).mp4 (3.4 MB, 800x1536)
3.4 MB
3.4 MB MP4
sometimes the hand movement is so fast not even genning at 2mp can fix it.
>>
>krea2
>wearing Itsuki Nakano cosplay
holy kino
>>
t2v pussy/asshole lora?
>bro just queue up 7 reference images and then you can
fuck off nigger
>>
File: 1780174026380798.png (408 KB, 1826x723)
408 KB PNG
Jesus
https://xcancel.com/bdsqlsz/status/2086034269940666788#m
>>
>>109497222
if they truly make it between 100B-200B like they said and train enough on it it should do so just by being big enough to remember it all
>>
>>109497266
that was just someone guessing btw, not a source from anyone who worked on it. No one knows. It for sure is not as big as sora 2. Could be like 50B
>>
>>109497262
just do
>tight white g-string panties wedged
and then use your imagination
>>
>>109497266
we unironically need bigger models. the amount of low IQ obnoxious jeets and brownoids in these threads has been a real issue. need to filter the scum out
>>
File: 1780122032032393.png (453 KB, 1080x985)
453 KB PNG
>>109497276
that anon is right, just stack more layers, who cares about finding a better architecture or improving on the training process!!
>>
>>109497266
KJ would find a way to make it work
>>
>>109497294
its likely a moe. In that case it would be optimized for running on multiple gpus at once
>>
File: 00184-3171973105.jpg (290 KB, 2880x1856)
290 KB JPG
krea2 is just the greatest blessing this year.
>>
>>109496626
the audio and lip syncing was poor before 2.3, but after was far better. Lip syncing could fail when you got that slow zoom in effect and basically that was just a failed gen and sometimes it just would not work with some prompts and/or source images. Overall that is an area you feel more confident in H3, usually it is a bad prompt when the wrong people talk are they don't talk etc.
>>
File: MiniMax_H3NoAudio_00019_.mp4 (2.23 MB, 1120x832)
2.23 MB
2.23 MB MP4
Where videos
>>
>>109497294
I think he released 4bit e3 this morning? I kinda skipped through the announcement as it didnt apply to me as im a medvramlet
>>
>>109497302
>>109497239
>>
File: 473444.webm (3.9 MB, 832x1184)
3.9 MB
3.9 MB WEBM
>>109497303
i think that is one of the sneedance shills that constantly lies about ltx in this general. pure text to video: https://litter.catbox.moe/a77ncy.webm
h3 can't do audio quality like this, or natural camera movements, or even the general aesthetic itself. h3 made a modern looking music video when i prompted for an 80s one
>>
>>109497311
doroguy here

my agent is extremely busy with tasks and story boards, not doro but she will come later for sure
>>
>>109497311
here
>>
>>109497311
I'm playing video games. Can't gen and play at the same time
>>
>>109497352
How did you get this video of me
>>
h3 keeps giving me unprompted nipples...
>>
>>109497343
tbf you can use ref2v and use images which are used to base the video on along side the audio and H3 probably is going to be more creative because it can simply do more with regard to physics. I think LTX should be viewed as what it could do rather than what it couldn't, there is little point in crying about something if it simply can't do it, you use it is within it's limitations.
>>
>>109497266
i doubt it, maybe if its moe
>>
>bro just spend an hour to find the right references, do the spaghetti shit in the workflow, write a prompt longer than a fucking book, and then
it's all so tiresome
>>
>>109497287
people are trying to improve the arch already, but more layers are always good to push for also since they bruteforce quality while allowing you to distill the models into smaller ones after
>>
>>109497389
I enjoy the process
>>
>>109497393
if you're doing actual art then of course
I thought we're trying to fucking jerk off here
>>
>>109497389
>bro just spend an hour to find the right references
what were you doing your whole life let alone the last few years if not downloading images/videos you like to use as inspiration/references?
>do the spaghetti shit in the workflow
mostly just works, ask llm after pointing it to the docs if you are retarded
>write a prompt longer than a fucking book
mostly just works, ask llm after pointing it to the docs if you are retarded
>>
>>109496435
ok but can it do dungeon elf grope scenes out of the box like I’ve done with h3?
>>
File: face_PNG.png (674 KB, 689x937)
674 KB PNG
>>109497389
Did someone say spaghetti?
>>
>>109497425
why are Black men so irrisistable?
>>
>>109497352
>>
>>109497373
if the model can do all of this just from text alone, then it should mean that the model has a better understanding of implicit things. prompting for an 80s video in ltx does not require you to start listing out technical terms for the things that cause analog film to look the way that it does. and funnily enough, it still can't do it from text alone which is why everyone is coping with references. why can't it do it from text alone? the dataset is slopped with CGI and video game footage in order to give the illusion of prompt adherence
>>
bloody benchod hours
>>
>>109497471
couldn't you just prompt for film like other models?
unless you mean shit like scanlines
>>
>>109496290
>car with hair
OP posting ugly shit as usual.
>>
>>109497493
>t. balding car
>>
>>109497394
you jerk off from the process
>>
>>109497491
maybe there's some secret keyword in there that will make it happen, but i tried explicitly prompting for vhs distortion and other terms that show up in analog film like bloom and film grain. doesn't do anything
>>
>>109497471
I would never rely on text only for genning unless it is for something that doesn't have a source image that can be generated at the very least or simply provided via an image already existing. I mean you could go and and say LTX can do things Wan 2.2 can't do (putting aside longer gens and audio), but the same can be said that Wan is superior in a lot of things. the models work differently and LTX simply cannot do some things Wan can do
>>
>>109497266
pretty crazy that a model 10 times smaller can deliver big time then
>>
>You could also add the RTX Video super Resolution node between the image input from Create Video and the VAE Decode Image output. only 20 seconds in addition for two time upscale!
https://github.com/Comfy-Org/Nvidia_RTX_Nodes_ComfyUI

worth?
>>
>>109497513
it's a nice experience to not need to keep unloading my video model to go load up an image model if i want to change something about the subject. the highly variable generations are also nice since the entire scene will be arranged differently each time. more changes for big surprises. ltx is smart enough to understand character references too if you desire that kind of thing as well
>>
>>109497544
i dont know
>>
File: 00194-1980104162.png (2.31 MB, 1920x1216)
2.31 MB PNG
https://civitai.red/models/2842933/aranea-highwing-final-fantasy-xv-krea-2-lora?modelVersionId=3209493
>>
>>109497544
maybe
>>
sup bros
Which model would you recommend for generating medieval fantasy pics (this is for my D&D campaign).

I have only generating slop porn so far
>>
>>109497544
yes, makes stuff look nicer at lower res
>>
>>109497544
I tried it, pretty lame. It fucks up faces that look fine at 0.2-0.3 mp. So maybe worth it if your original gen was already high-res
>>
File: enough.jpg (40 KB, 403x392)
40 KB JPG
bros I'm thinking of raiding the nearest data centre to steal some RAM and GPU
>>
>>109497572
You get the death penalty for that
>>
>>109497544
it's like all upscalers, garbage in garbage out.
more wiggle room for anime and cartoons, but if you are trying to upscale realistic gens they need to be good from the get go.
>>
>>109497572
You wouldn't even be able to run the kinds of gpus used in data centers. You'd need direct access to the power grid
>>
>>109497544
it doesn't fix fucked up small faces
>>
https://huggingface.co/Kijai/MiniMax-H3-experimental
verdict on this?
>>
File: 1762289604621524.mp4 (3.33 MB, 1056x608)
3.33 MB
3.33 MB MP4
2 clips, 8 steps in post, 20 in link, minor rule infringments and 0.6mp text to video rugby clip , 8 steps really struggled with the very long AI genned prompt which was 1,195 words.
https://i.4cdn.org/wsg/1786190681400715.mp4
>>
>>109497389
nah I've set up my shit so it just looks at images and turns them into porn automatically.
>>
>>109497572
just get a used 3090
or are you so poor you can’t even afford that?
>>
>>109497565
krea has insane amounts of licensed material & weebshit knowledge and you can go for realism too
>>
>>109497343
>vid
Is the one on the top here H3? or Flux
>>
Sigma shift feels like voodoo to me. Can someone explain what it even does in the context of H3 output and what I should be looking for? Combined with seed variance it's hard to tell how it matters.
>>
>complains about H3 limitations
>'just use a reference'
>waah but I want text only
Are you shills for real? Is that how you'll advertise flux 3 lol?
>>
>>109496488
no, his loras are excellent, in top 1% ever made and published
>>
>>109497642
I don't know what I'm talking about here but I think lower video shift makes it spaz out with faster movement (good for high action), while higher video shift makes it slower but more detailed and higher quality.
>>
>>109497638
top is h3. bottom is ltx
>>
>>109497662
i know it's not what you prompted for but i like the top one better lmao
perhaps it's overtrained on modern shit
>>
>>109497657
If it's any use i used vid sigma shift of 12 for those two rugby clips i just posted and 10 for the audio, the audio difference was unnoticable between 8 and 20.
>>
>>109497674
12 is the balanced value. well, it's the one that everyone uses, so that makes it balanced. but i heard it's not the default value.
>>
>>109497645
>flux
the fuck does some jew shit model have to do with anything
the fact remains H3 can't even generate a fucking pussy. it's as simple as that.
>>
>>109497667
i should try a modern music video with ltx
>>
>>109497686
give it a reference, chuddie. it just works. and it's right out of the box. no lora required
>>
File: 122733CUI_00001_.png (588 KB, 1216x832)
588 KB PNG
>>
>>109497699
right out of the box and twice as slow
>>
>>109497710
with images, nope. unless you're retarded and feeding it a giant video
>>
File: 1757401019795581.jpg (228 KB, 2000x1333)
228 KB JPG
what is the use case of flux 3?
>>
>>109497651
top 0.1% most verbose captions you mean
>>
>>109497716
pressure on china
>>
>>109497716
i'll take any open source weights i can get. even if it can't do NSFW, it might have meme potential
worst case it just fades into obscurity
>>
>>109497716
flux 3 will be cucked and useless like all the other bfl models
>>
its not even worth discussing flux 3 if they wont do something like a klein distillation for it
>>
whys this debbie downer in here attempting to spin a negative light on everything
>>
>>109497716
to make astronauts riding horses on the moon and impressing redditors
>>
>heeeeeerr why das ist speaking ze truth??
>>
File: png.jpg (35 KB, 400x387)
35 KB JPG
https://www.dellstore.com/alienware-area-51-gaming-desktop-cadaat2265cto02mino.html
If I buy this how many years will it take me to make back the money I invested in this shit?
>>
I'm trying to do a gen of a character using the reference model that starts wide and zooms into their face. I figured I'd use a reference image for their face (that's from a completely different scene) at a closer angle as a second image input, however I"m finding it often times just treats that second image as a keyframe for the end of the video. Anyone have success with this? It doesn't happen every time. I have the Subject tied to the reference image and the Subject is listed as partially_preserved, I also tried attribute transfer and weak reference but that second image still seems to be treated as a keyframe like half the time. I even tried putting several different angles in one image. Maybe I could try rembg so there's no other content in the reference image at all.
>>
File: alien-cat-alien-dog.gif (701 KB, 165x124)
701 KB GIF
bro got that goon goon ging gang ayy lmao computer
>>
>>109497782
Depends on how much you value local genning. The market is so fucked that prebuilts sometimes have great deals on them occasionally since they bundle in so much other stuff.
However, the shame from buying alienware crapboxes will never go away
>>
>>109497782
after buying an apple product and gay sex (in that order), the next gayest thing you can do i buy anything alienware.
>>
what time does it take you guys to gen 20 steps 1mp 10seconds no references no nothing
>>
>>109497824
just over half an hour according to my notes
>>
>>109497782
>₹858,890.00
what the absolute fuck is going on with prices of pc in india? is there some tax bullshit the government placed on pc components sold abroad?
>>
>>109497840
lol the US equivalent is $7k while that one is $9k
>>
>>109497302
it would not be as good without you
>>
File: file.png (26 KB, 1419x354)
26 KB PNG
>>
>>109497865
if only sol attention didn't increase vram usage
>>
>>109497865
bro your quality?
>>
File deleted.
>
>>
>>109497782
your izzat will never recover >>109497791
ayy lmao
>>
>>109497892
>>109497728
>>
File: 1769464969487684.gif (1.13 MB, 498x498)
1.13 MB GIF
>>109497892
>>
>>109497899
This but remade with H3
>>
>>109497865
>9:11 at 10 seconds
>14:40 at 15 seconds
Good to prove wrong that dumbass who keeps saying "gen times get exponentially higher!"
It's linear and he's just spilling his shit over into normal RAM.
>>
What do you guys use to rewrite your prompts again
>>
use qwen3vl
avoid gemma 3 because it cant <think>
>>
>>109497914
>gemma 3
bro, your gemma 4?
>>
>>109497906
Global SOTA omnimodal model (White) man's brain.
>>
>>109497906
there was a learning curve but I think I've got the syntax down now. Haven't tried everything yet but I've got it to do some neat things I didn't think it would get right
don't need assistance any more
>>
this video character replacement stuff is really hit and miss
I don't think you can write a generic prompt for it
>>
>>109497824
10.5 minutes with the cope optimizations. 16vram/64ram.
>>
>>109497941
i hate that i've seen this clip of the dancing russian kid so much that you just replaced him with a huge miku plush and i still knew it was that clip
the chick in the background is fucking insanely hot however
>>
>>109497904
i believe it does increase exponentially when u increase the length and u ran out of VRAM
>>
this nigga talm bout exponentially quadratically increasing gen times
>>
>>109497948
not as bad as that one anon recognizing which porn someone was referencing
>>
does anyone actually use the prompt enhancement shit built into the krea2 workflow
>>
>>109497343
>>
>>109497941
Brother if you follow the official syntax the ref model prompts are like a 4 page document, front and back.
>>
you guys remember when prompt shops were a thing?
>>
ref model use more vram? I can gen in 15 sec 1.5 mp, the latent pass only noise till I low the resolution or reduce seconds...
>>
Has Furk given his thoughts on H3 yet
>>
h3 patch sage and mem eff nodes dont seem to do anything if you already use the sage cli arg
>>
>>109498009
Subscribe to his patreon to see it, including the most efficient workflows for it!
>>
>>109498009
Now that's a name i haven't heard in a long time.
>>
File: file.png (1.1 MB, 1530x1393)
1.1 MB PNG
>>109497983
I know
>>
>>109497946
20 steps??
>>
File: booba.gif (2.06 MB, 498x395)
2.06 MB GIF
>>109498043
post that
>>
>>109497622
>implying used 3090's are cheap
>>
>>109498043
POST WITH METADATA PLS
i think i've just popped my first funny boner in these threads ever, looks great and awful at the same time and also hot.

ahh wait my boner just deflated realizing that's the level of prompting i'll have to get to to make anything that good..
>>
>>109498058
>ahh wait my boner just deflated realizing that's the level of prompting i'll have to get to to make anything that good..
looks like a bunch of nonsense that an llm generated which doesnt do anything
>>
>>109498067
don't go prompt coping too, pal. it's pretty obvious at this point the reason most of our gens in this thread stink shit is because they're not detailed like that.
if it's that easy for any LLM to spit something like that out then the skill ceiling isn't that high.
it's just laziness at this point.
>>
>>109498048
I guess spectrum simply skips half the steps
>>
>>109498048
Yes picrel is log for this >>109497973
>>
>>109498050
because the original is blurry noisy the ouput is too
https://files.catbox.moe/1f57e2.mp4
>>
File: file.png (202 KB, 600x496)
202 KB PNG
Why i get this? Someone here had this problem with ref model in 1.5 mp 15 sec?
>>
>>109498081
>1 second

cmon son
>>
>>109498084
I want to stay what is under the original gif when testing
>>
>>109498057
are you saying you don't have between $1000 and $1500 available to you?
it's still the best price for performance ratio you can find for this hobby
>>
>>109498081
>stiff, unmoving breasts
as expected of H3
>>
>Get oom at 30 seconds at 1mp
>half way through the 10 minute generation
I hate this shit......
>>
>>109498105
>best price for performance ratio you can find for this hobby
If you're going purely off price for performance ratio a 16GB 5060ti is way better value
>>
>>109498105
It feels like such a scam that card is going for that much. I really fucking wish AMD or intel will get it together because if they did all of these cards would be cheaper and most of us would be rocking multi RX pro builds.
>>
>>109498117
I had a 40min gen that failed at the save video stage for some reason...
>>
>>109498135
H3 looked at your output and was like "oh hell no I'll do some degenerate stuff but this is too much"
>>
File: MiniMax_H3_00568.mp4 (2.54 MB, 1440x1088)
2.54 MB
2.54 MB MP4
>>
>>109498135
That fucking sucks.....
When using a reference the model can typically keep cohesion at 30 seconds so now I'm testing if it can when doing text to video only, I'll post the result if it's not shit
>>
>>109498148
trooned malfoy is cute
>>
>prompting ltx for better breast jiggle results in worse, choppy, air balloon kind of jiggle
>>
>>109498113
they only become stiff in the club
when she gets put on beach they are more jiggly
>>
>>109498163
meant H3, never had such problem with ltx
>>
Am I the only one who thought will smith eating spaghetti was cringe? How can people still be clapping like seals at this unoriginal shit?
>>
Could someone catbox a good-looking anime-styled video made with the reference model so I can check what nodes I can use without dropping the quality too much?
>>
>>109498172
everyone except one guy (who i also now suspect is a janitor) finds it cringe.
>>
>>109498172
because of the effect of the eternal summer, the normgroids learn a simple, old, "in-group" meme but it never stops being spammed because the hobby has a constant and ever increasing influx of newfags.
>>
>>109498172
yes it's tired, anyone still using it knows this and I assume they're doing it in malice
>>
>>109498172
I’m new to vidgen and find them funny, but you fags constantly bitching about them are pathetic.
reminds me of nintendoschizos on /v/ mindbroken by some eric
>>
(adult woman: -1)
its genning time
>>
FRESH
>>109498215
>>109498215
>>109498215
>>109498215
>>
>>109498199
>generates adult man
>>
>>109498216
No thanks faggot, you really need to use your time better instead of posting mentally ill gens in your containment and saying GM to the same 3 retards which may or may not be you coping
>>
File: pepe-shrug.png (194 KB, 508x492)
194 KB PNG
>>109498221
wat?
>>
>>109498232
You spammed the change OP move too often retard, fuck off.
>>
>>109498177
I'm just using the default workflow though
>>
>>109498249
what is your problem? i'm just copying the relevant information into the new thread
>>
>>
do not use the troll thread. someone is baking a real thread
>>
based
>>
Someone got a good upscaler for anime video? The ones I used for imagegen aren't doing great
>>
>>109498253
Without any extra nodes? Could you post a good output and the prompt?
>>
For anons using sigma shift, does that parameter need to be seen by both the guider and the scheduler or is it correct to only feed that output to the guider?
>>
>>109497865
tried adding sigma spectrum and sage
yes it's faster but the output is chopped
>>
>>109498216
change the fucking name of your threads because know one here wants to be associated with it.
>>109498261
because he is a fucking troll, some of us are adults and can't be bothered with silly fucking games.
>>
>>109498371
>because he is a fucking troll
how?
>>
Is anyone going to make the non-debo thread? This autistic loser never seems to sleep
>>
If my video gen completes I will make a new thread with that image. This schizo dude is really sad and he will still spam links he won't vet for malware because he can't help himself.
>>109498394
While true he has friends that help him all equally fucked.
>>
>>109498394
other baker got a 3 day ban for being too based
>>
new
>>109498215
>>109498215
>>
>>109496779
Those guys went closed after the first Kolors released, but who knows, if Minimax changed their minds, they could as well
>>
>>109496435
>Gemini Omni is better than Seedance 2.0 in the chart



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.