[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology


Thread archived.
You cannot reply anymore.


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109499734 >>109500977

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

08/08/2026

>Kijai: MiniMax H3 Ref Lora Rank 256 bf16
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras

>MiniMax H3 at native fp16 on pre-bf16 GPUs (V100 / Volta)
https://github.com/Amduraznak/minimax-h3-fp16-fix

>Cosmos3-Nano-WebUI: Self-hostable API + Web UI for Cosmos3-Nano quantized fp8 and nvfp4 checopoints
https://github.com/fengwang/Cosmos3-Nano-WebUI

>R9700 AI Pro — ComfyUI / MiniMax-H3 speed patches
https://github.com/charlie12345/R9700AIProComfyUIPatch

>MiniMax-H3-Pruned-GGUF
https://huggingface.co/Abiray/MiniMax-H3-Pruned-GGUF

08/07/2026

>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely local
https://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha

>LIGHTX2V 4-step Turbo Minimax H3 lora
https://huggingface.co/lightx2v/Minimax-h3-Turbo

>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRA
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>Sage Ready: Local-only installer and readiness checker for SageAttention
https://github.com/CosmicFungi/Sage-Ready

>Wan 2.2 Animate 2 14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B

>MiniMax-H3 FL2VA — MLX-Serve, 2-bit text encoder / 4-bit DiT
https://huggingface.co/antocorr/MiniMax-H3-FL2VA-MLX-Serve-2bit-text-encoder

>H3 Motion Context: Clip chaining for MiniMax H3 in ComfyUI
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

>ComfyUI MiniMax H3 FirstBlockCache
https://github.com/duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache

>KVAE: Family of Tokenizers for Multimodal Generative Models
https://github.com/kandinskylab/kvae

>Energy-Guided Flow Matching
https://github.com/ysng123/EG-FM

>VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing
https://zzzmyyzeng.github.io/VideoArgus

08/06/2026

>Flash-VAED: Plug-and-Play VAE Decoders for Efficient VidGen
https://github.com/Aoko955/Flash-VAED

>(preview) MiniMax-H3 Turbo LoRA
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora
>>
>nigbo malware
>>
>>109502332
>not having an LLM to handle the prompt in real time
>>
>mfw Research news

08/08/2026

>Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation
https://arxiv.org/abs/2608.04902

>Coherence-Oriented Dream Scene Visualisation
https://arxiv.org/abs/2608.05233

>Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation
https://arxiv.org/abs/2608.05210

>GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
https://arxiv.org/abs/2608.03083

>IRIS: Visual-Semantic Binding for Forgery-Resistant Watermarking of Diffusion Images
https://arxiv.org/abs/2608.03539

>Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
https://arxiv.org/abs/2608.00663

>A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
https://arxiv.org/abs/2608.05260

>Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
https://zhangbo135.github.io/EviSelect

>Progressive Decision-Making for Localizing Open-Ended AI-Generated Image Forgeries
https://arxiv.org/abs/2607.29156

>GeoMAR: Unleashing Geometrically Aligned Features for Masked Autoregressive Blind Face Restoration
https://arxiv.org/abs/2608.03923

>WorldClaw: Agentic 3D Open-World Generation at Scale
https://arxiv.org/abs/2608.05248

>UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space
https://arxiv.org/abs/2608.03817

>Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference
https://arxiv.org/abs/2608.03867

>Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models
https://arxiv.org/abs/2608.03160

>Attention is Case-Sensitive
https://arxiv.org/abs/2608.03711

>In-Context Collapse in Vision-Language Models and How to Mitigate it?
https://arxiv.org/abs/2608.02830
>>
>>109502333
blessed triplets of frens
>>
>>109502333
thanks for the bake
>>
hmm, now everyone is using gemma for prompting minimax h3 while no one was doing it for ideogram despite similar rigid prompt formatting requirements
interesting isn't it?
>>
>>109502354
>how is he so fucking fast with his garbage news?
probably a bot
>>
breast thred of goonship
>>
>>109502335
i do not wish to use a cloud service
>>109502334
way ahead of you but what about the other one? im horrible at describing abstract shit
>>
>>109502367
ideogram kept saying what I was doing was haram and it kept sending my location to qatar
>>
The model is capable of doing really crazy stuff with the camera, full POV and all that.

Now I'm wondering has anyone tried generating (or continuing) an SBS video with H3 yet? I'll be doing it tomorrow for science.
>>
>>109502374
Use gemma locally for that shit
>>
>>109502367
It's just an effort to reward ratio, and the main problem with ideogram was the cuckboxes you were forced to draw for every gen
>>
>>109502378
I'm going to generate an H3H3 video with H3.
>>
>>109502377
>minimum number of bboxes: 5
just add this to your prompt generator instruction and it generates anything without triggering filter
>>
>>109502383
any node recommendations for this so i dont have to install another llm toolkit or anything? i wont ask you for specifically amd solutions but ideally if you have one that would be handy
>>
>>109502390
ideogram doesn't generate hot trannies in my area. not interested.
>>
https://files.catbox.moe/lreg9t.mp4

good shit
>>
>>109502333
Damn, it definitely helps though. Like night and day difference. I'll try it with the Turbo LoRA workflow now

Baseline
https://files.catbox.moe/8rzon3.mp4

Fix
https://files.catbox.moe/128mbo.mp4

Fix, slo-mo
https://files.catbox.moe/a0frh1.mp4
>>
>>109502398
real shit fr fr no cap
>>
>tranime
I'm out
>>
>>109502398
oh my god i've not heard that audio in ages
>>
>>109502398
>>
>>109502402
Mean to quote
>>109502328
>>
>>109502402
That seems smart. Gen bunch of shit and second pass only the stuff you like
>>
So /ldg/, are you guys more of a [keyframe completion] community? or a [reference generation] community?
>>
File: oh well...gif (2.23 MB, 279x374)
2.23 MB GIF
>>109502402
>>109502417
>2 mn to 20 mn
>>
damn ok. turbo lora is fucking retarded with ref model. good to know.
>>
>>109502408
https://files.catbox.moe/rda7zi.mp4

feast your ears
>>
>>109502432
>model drops
>takes a long enough time to make one gen even on high end that people already make several turbo loras
>several custom nodes to speed it up as well including various attention patches
>just to make them gen faster at the cost of quality
>the one thing to fix the quality also fucktouples the gen time
>we will soon see workflows so fucked up from the speed that the base gen will be 20 seconds and the second pass will be 30 minutes
>>
with the ref model being so good with high-res sources, we might actually have a workflow somewhere in the future where the gens are extended beyond 15s accurately in a single run and stitched.

I'd try to make one myself but I'm lazy and retarded
>>
File: MiniMax_H3_00363_.mp4 (2.82 MB, 832x640)
2.82 MB
2.82 MB MP4
So this is the power of the turbo lora.
>>
File: tomb of wan the 2nd.mp4 (3.55 MB, 1216x672)
3.55 MB
3.55 MB MP4
>>109502438
this is what glimpsing the infinite feels like
>>
File: 35746.png (27 KB, 1067x224)
27 KB PNG
what did bro do?
>>
>>109502448
>I'd try to make one myself but I'm lazy and retarded
you already made it just by having the idea, now paste your own 4chan post into an LLM
>>
>>109502435
Works for me, use 600ema version, 8 steps euler/beta
>>
>>109502450
kino. what the hell
>>
>>109502456
see >>109502449
I used those exact settings. just made it really retarded. The output looks really nice for 0.5mp tho. no artifacts.
>>
Anyone tried https://github.com/matlowai/ComfyUI-MAINodes?
Good or bad for face for ants syndrome?
>>
>>109502450
>not the millionth shitty meme about fent, floyd, niggers or jews
KINO
>>
File: should be all right.jpg (62 KB, 1081x558)
62 KB JPG
im not getting good results at all with the new 600 turbo lora, everything is all artifacty and grainy. doing 0.5mp. did i set this up wrong somehow?
>>
>>109502470
Purchase an advertisement
>>
>>109502484
Reduce strenght to 0.9
>>
once again trying cope sol-attn
>>
>>109502461
I think prompting matters too, I had bad quality gens on H3 when the prompt isn't coherent, similar to LTX bad experiences when prompting something the model/text encoder doesn't understand right, and sampling is fighting with itself turning the results into mushy garbage
>>
File: screenshot.1786230377.jpg (364 KB, 1497x822)
364 KB JPG
syntax highlighting added for MY prompt enhancer ;)
>>
>>109502495
well, it's better than spectrum, but... still shit.
>>
>>109502485
Get some publicity material
>>
>>109502499
*claude prompt enhancer
>>
>>109502498
I agree that the prompt is important. This isn't a prompting issue tho, the gen was fine before trying it with turbo.
>>
>>109502501
>it still causes morphing.
why do people use this shit?
>>
>>109502484
remove external turbo sampler node, select euler
that node is ass, designed for 4steps
>>
>>109502527
to stop coping one needs rtx blackwell 6000 pro.
EasyCache, spectrum, sol-attn, turbo lora - everything's shit.
>>
>>109502538
I don't have that amount of disposable income.
>>
>>109502540
sir, comfy cloud is only $20 away
>>
>>109502538
The only two optimizations that aren't too insane :
- sage attention 2, no perceptible change from my tests, and nice boost
- int8 convrot model (same, no change for h3) using the pruned model
>>
>24gb vramlets coping
use comfycloud to access enterprise hardware for cheap. vidgen isn't mean for your shitty 2022-era hardware
>>
>>109502540
Then pay with your time
>>
>>109502538
I'm considering it, but very reluctantly. They can't get more expensive than the current $13k, right?
>>
>>109502544
>>109502541
and how I should going gen my cunnies using copper cloud sirs?
>>
>>109502549
holy shit, that's what they're at now? i remember when they were $8k
>>
>>109502538
So how quick does it gen a 1mp 10s vid?
>>
>>109502549
They probably will if nvidia next release cycle is to be trusted :
- blackwell gaming "SUPER" refresh for early 2027
- actual next gen 2028

-> means these 96GB cards will be in demand at least until then.
>>
>>109502551
fearlessly
>>
>>109502549
You mean barely $20K at the end of the year
>>
at least sol doesn't seem to fuck with prompt adherence
>>
>>109502549
at my place they already cost like 20 grand
>>
>>109502499
You can do the exact same thing with LM Studio. And it's got way more features.
>>
I've been using Qwen3.5 9b to generate prompts for H3. How does it compare to gemma? Worth downloading gemma?
>>
>>109502571
gemma is more creative
>>
Buying a GPU right now is retarded. In a year or two, the AI bubble will burst, GPU prices will come down and Nvidia will shift its focus back toward the consumer market. We'll probably start seeing enthusiast AI hardware with much larger amounts of VRAM. I'm talking things like 64GB for around $1k
>>
>>109502571
Qwen sucks. Gemma is better. I get best results with 26B A4B heretic.
>>
>>109502537
still came out just as grainy, fuak
guess i'll just wait for a better turbo to come out later
>>
>>109502574
lol
>>
File: MiniMax_H3_00365_.mp4 (2.92 MB, 832x640)
2.92 MB
2.92 MB MP4
>>109502449
for reference, same seed
>no turbo
>sol-attn, err_sde / beta57 15step
>0.5mp
>395sec (6min35sec)
>rtx 3090
>>
That's good bait
>>
I remember my first Turbo lora
>>
>>109502574
this is a extremely dangerous level of copium to ingest. you should consult a medical professional immediately
>>
>>109502432
>>109502441
Not so fast
https://matlowai.github.io/ComfyUI-MAINodes/#ladder

That middle gen only took 2.2 mins for the 2nd pass and looks better than raw Turbo and raw baseline
>>
>>109502574
i listened to people like this about ram, and look where that got me
>>
>>109502574
>I'm talking things like 64GB for around $1k
This should have happened long, long years ago. Dram was cheap as dirt yet nvidia only supplied 8GB for gamers, and that in the age of quickly rising 4K, forcing game devs to use retarted techniques to have the textures looking okay.
>>
>>109502578
You need 22k in computers to make a good prompt
>>
File: 1541461395847.png (180 KB, 300x318)
180 KB PNG
The reference model is genuinely magic. I can't believe some of the shit I've managed to do with it via motion transfer. Literal holodeck tier shit.
>>
>>109502574
yeah yeah, bubble will burst, we will find our tech jobs again, gpus and ram will cost less.
simply won't haplpen.
>>
>>109502589
The sweet...sweet turbo lora...I ALWAYS HATE IT IT!
>>
>>109502607
kek
>>
>>109502601
>deepsuck
kek
>>
>>109502604
Does it really do motion transfer without fucking things up?
I was only using i2v from the other model...
>>
>>109502574
here's a tip: if you have to fantasize about how you'd be living in a splendorous perfect a future once the bubble bursts, then there is no bubble. this is like saying
>i cant wait for the inflation bubble to burst so i can be rich again soon!
>>
I wish H3 had that magical second pass LTX had that mad the output look perfect.
>>
I am what the late /r/ called a "wizard".

You are merely a boy.
>>
>>109502591
>>109502625
Just look at the trajectory of where AI is going. As local AI continues to improve, reliance on cloud services will decrease, making many of those services increasingly difficult to sustain. MiniMax has already made a huge dent in the video genning space, while China has also made a significant impact in the LLM space.
From the beginning, China's strategy has been to disrupt the American cloud AI business, and that strategy is becoming more effective as its models improve. The bubble bursting is inevitable, as local AI continues to close the gap with cloud services
>>
>>109502623
nta but if the prompt is good and the source clip/positions generally match whatever you're transferring it onto, yeah
>>
>>109502398
>https://files.catbox.moe/rda7zi.mp4
>someone finally used this audio i linked many threads back
SOME GOOD SHIT RIGHT THERE IF I DO SAY SO MYSELF I DO SAY SO
>>
>>109502645
It apparently has it, but minimax didn't release it yet. Some kind of upscaler to reach 2k from the advertised 720p of the model.
>>
if the AI bubble will burst (it should, there's no way to sustain its unprofitability) it will sweep the entire world's economy, and we'll have a great reset or something like that
>>
>>109502574
give me your copium and hopium dealer's name, I need some too
>>
File: 1769708309312401.png (161 KB, 1735x1483)
161 KB PNG
>>109502402
Apparently it's fucked (checked with gpt pro model).
>>
>>109502664
Yup, and it's happening in exactly 2 weeks.
>>
>>109502657
lol. you realize that you have to pay more than a cloud subscription if you want to run competent AI models locally, right?
>but muh stable diffusion and sageattention!
nobody cares about image/video, it's a nothingburger. it doesn't matter if deepseek beats gpt or claude, you still need millions of dollars in hardware to run it at full potential which is why local consooomers will never have affordable vram. openai could crash tomorrow and there would still be millions of B-list companies lined up to buy their hardware before you ever get a chance at it.
>>
Pulled latest comfy and spectrum and now there's like a weird desync between the in-comfy progress and sampler preview and the step count outputted in the terminal. Like, in-comfy's displays are lagging behind the terminal steps, and the default sampler preview is all stuttery and fucked, or doesn't play.
>>
File: hooooooooooOOOO.jpg (410 KB, 1504x2664)
410 KB JPG
>>109502398
>>109502408
>>109502438
OH YEAH here's the video i was talking about that someone did in ltx. full 20 seconds of coherent super high quality video one could even say it's some GOOD SHIT RIGHT THERE MMHHMMHHMM

https://civitai.red/images/130079342

https://files.catbox.moe/0uptmb.mp4
>>
>>109502679
>spectrum
delete this cope shit already
>>
>>109502574
The US dollar is probably going to be Zimbabwe-tier in a couple of years. A loaf of bread will be $1k dollars.
>>
>>109502664
Every other day someone predicts the AI bubble burst, and it's still going since at least 2023, with more investment than ever in robotics, various models, and especially in memory (from colleagues, there is a gigantic scramble to get new production units running to make this by 2027).
>>
>he pulled
>>
>>109502367
But I did?
>>
>>109502367
video gen is more interesting than image gen
>>
>>109502574
>>109502688
Kek you people are delusional.
>>
>>109502690
Yeah, they're burning your retirement money.
There is no profit. The money will run out
>>
>>109502574
You realize we are in the niche, right? Normalfags who use AI pay for SaaS/cloudshit. Local is utterly irrelevant, if it wasn't, companies would never open-source useful models, but since there are so few people capable of running them at acceptable speeds, they don't really lose their margin of profits, it's no big deal and they also give others the opportunity to improve their software stack. I genuinely think the only reason OpenAI, Google, Anthropic and others don't release weights is because they don't want competitors to investigate how they trained their models and got them so good, other than that no one but large datacenters are able to run their models, so I doubt they are concerned that a rare rich autist wouldn't pay for their API because he can run the weights local (I bet almost no one in /lmg/, even the richfags, runs Kimi K3 regularly at home with full precision).
And also, if selling retail GPUs somehow cannibalizes the still-profitable datacenter business, Nvidia and the others will not give a fuck about us
>>
>>109502574
Do not listen to him. If you don't buy your hardware now, you'll regret it in a few months when things will be even more expensive. No one wants you to own things, not the gov, not openai or nvidia. They want to tie you to a cloud subscription and making things expensive help them achieve that goal.
>>
File: MiniMax_H3_00368_.mp4 (2.3 MB, 832x640)
2.3 MB
2.3 MB MP4
>sol + cache
>res_multi / simple 25steps
>5min20sec
god I love cope nodes. the quality is impeccable
>>
File: 1779094392571544.mp4 (2.42 MB, 1216x672)
2.42 MB
2.42 MB MP4
i remuxed all my files to try and save space and in the process lost all the prompts so i just deleted them which was probably healthier
>>
>>109502574
At this point I’m pretty sure this is just what things cost now and your job simply makes them seem expensive.
>>
Someone make a ps2 graphics lora for Anima
>>
>>109502667
Yeah I've no idea, just tried it with their Turbo wf (which is just Turbo 2nd pass, first gen is regular) on a different prompt and I'm getting smearing so I'd have to use the normal model to see its benefits
>>
File: 1760915949477945.png (8 KB, 1656x69)
8 KB PNG
You said that int8-convrot vae is good, but vmaf says there's a difference between fp16 and in8-convrot, and quite a big one (vmaf).
>>
>>109502709
This gen is missing Will Smith with all those spaghetti on the floor
>>
>>109502722
I tested it a few times myself and the speed savings were barely more than the fp16
There was weird smearing and artifacts too so I decided to leave it alone and work on other ways to gen faster
>>
>>109502702
Sure anon, sure.
>>
>summary:
>[video continuation]

>detailed_description:
>ahh ahh mistress
>>
>>109502703
I'd say we are in the same position as the game consolefags who are complaining that Sony is abandoning physical media and going all-digital. They probably looked at reports, determined that most people buy games digitally, knows that selling digital copies is more profitable than allowing people to use second-hand physical media, and decided they would go all-digital.
Running AI on gaming GPUs is similarly a drop in the ocean when compared to cloud usage, so there is no reason to pander to us (yet)
>>
ok. hear me out.
>0.4mp first pass
>upscale to 0.8mp
>4steps turbo lora at low denoise.
>>
>>109502736
Well, I disabled sage-attention to test it. Two videos, same seed. There's a, like, 0.2 difference in vmaf between two videos in common.
But between int8 vae and fp16...
Don't use it anons
>>
https://old.reddit.com/r/StableDiffusion/comments/1vjaq5e/minimax_h3_pinkcherry/
>>
Buying a house right now is retarded. In a year or two, the real estate bubble will burst, house prices will come down and banks will shift their focus back toward the middle class market. We'll probably start seeing big houses with much larger amounts of land. I'm talking like 64 acres for 100k.
>>
whats the veredict on low step loras?
>>
>>109502709
Awesome
>>
>>109502760
thanking you for contribution sar
>>
File: ai reacts to me cum.jpg (593 KB, 1851x1240)
593 KB JPG
>>109502747
kek, good times
>>
>>109502763
Scheisse
>>
>>109502679
I got that problem earlier too, anyway i switched to the turbo lora
https://files.catbox.moe/kq0nw8.mp4
>>
>>109502763
>whats the veredict
video okay saar
but audio very bad
>>
BLAME is cool but has anyone done any Berserk animations yet?
>>
>>109502768
>again, no offense
>>
>>109502780
Berserk already his plenty of real animations tho. BLAME! only has that one shitty netflix movie.
>>
>>109502773
audio working fine on the ema600 checkpoint and no shift node either
>>
>>109502788
Nothing past Golden Age has good animation though.
>>
>>109502771
>vram debug node
won't it load models from the disk each time the generation starts? my nvme is fast, but still raping disk isn't good
>>
>>109502788
>adapting the entire thing in the 97s style
holy kinoli

fr though, it's fucking depressing that such a beloved anime got a bullfuck adaptation, meanwhile absolute slop gets a budget
>>
>>109502810
Was testing out if it solved the hitching issue at the start of gens. It did actually decrease the overall time
>>
H3's days are numbered
>>>/wsg/6210574
>>
https://www.reddit.com/r/StableDiffusion/comments/1vj3l5b/minimax_h3_spectrum_v021_new_offline_replay/

updated
>>
>>109502800
well yeah, notice I didn't say anything about the quality of those animations lol.

There's a bunch of reasons I'm going with BLAME!
Thematically it's already pretty closely related to AI generation. The art already lends itself really well to black&white (color matching would be hell). And Nihei's pretty loose with the designs of the characters. lots of differences from page to page which makes the character inconsistencies between gens kind of mirror that.
>>
>>109502402
>>109502716
Ah, featherweight stack workflow https://matlowai.github.io/ComfyUI-MAINodes/#featherweight, trying it right now
>>
>>109502821
Not reading that sloppa, if I wanted to read claude I'd just ask it
>>
File: rawrxd.mp4 (2.27 MB, 832x1248)
2.27 MB
2.27 MB MP4
Finally my GOONINATOR 3000 can have video slop
>>
>>109502821
Good job, it broke the gen preview in the sampler
>>
>>109502832
oh also, barely any dialogue.
>>
>>109502816
The movies were decent, especially the 3rd. Shame they never got the chance to continue them.
>>
>>109502821
qrd?
I fucking hate people who unironically paste that kind of sloppa
>>
>>109502856
Ask your AI waifu to summarize it.
>>
git: pulled
drivers: updated
its kino time
>>
>>109502821
>Yeah, I'll fix that in the next patch, just wanted to get the release out of the door asap
So you push it out with one of the most critical features, a fucking preview of the video, not working. Are you simple?
>>
>>109502835
>Speed: on our card the w4a8 pipeline ran somewhat slower per clip than int8 (29 vs low-20s minutes for the 5 s case); on a 32 GB card that is the wrong comparison, because int8 does not fit and offload-thrash costs far more.

Nvm, going to a different quantization than int8 is just dumb
>>
>>109502870
There is still hope

https://files.catbox.moe/p00k43.mp4

So far I've been using just sage. Maybe the time for improved generation can be cut in half with cache? https://files.catbox.moe/p00k43.mp4
>>
>>109502687
>delete this cope shit already
Spectrum turns a 13 minute gen into a 7 minute one for me, and the results are pretty much visually identical
>>
>>109502863
workflow: BROKEN
>>
>>109502712
Tempting. Would be even moreso if you provided a dataset.
>>
>>109502866
it's fixed as its at 22 now
>>
Anyone tested if nvfp4 text encoder really breaks coherency and overall way worse than int8?
>>
>>109502927
i'm on the latest and its still happening for me
>>
Where are the dance videos?
>>
I need like 20 5090s to get all the ideas out of my head.
>>
non_diegetic_music: N/A

yep, it's quiet time
>>
So does Sigma Shift + Spectrum fuck up handheld camera movement? I think it does for me. Like sometimes it will produce a vibrating camera.
>>
>>109502933
Haven't compared but it MAY make a difference for the ref2va model. In my tests, it is struggling with large prompts
>>
>>109502883
idk how people prefer cache over spectrum just look at thje two side by side on the right
>>
File: MiniMax_H3_00446__2.webm (1.99 MB, 864x608)
1.99 MB
1.99 MB WEBM
>>
>>109502941
bro I had to fire up the 5090, the 6000 alone wasn't enough. It's unreal the reference shit.
>>
File: 1780029311694912.png (5 KB, 180x180)
5 KB PNG
>>109502863
Isn't there someone you forgot to ask?
>>
>>109502953
Why is she squirming like a retard
>>
File: 3157140495.gif (1.13 MB, 498x498)
1.13 MB GIF
>>109502942
>integrated_multimodal_description: N/A
>overall_soundscape: N/A
>non_diegetic_music: N/A
>>
>>109502953
post new gens.
>>
>>109502952
speed nigga
besides you're probably gonna need to regen anyway so speed is always better
>>
>>109502954
It really is. It can be tricky to figure out exactly how to prompt for the thing you're trying to render, but once you get it figured out... holy fuck...
>>
>>109502954
>I had to take the Porsche out for a spin, the Ferrari just wasn't enough
meanwhile I am happy using my Ford Fiesta 1.4
>>
>>109502952
Biggest difference i see is a forward roll vs a side roll and she still gets ripped into two bodies in both
>>
>>109502402
Motion is not always better, but it does get rid of the smearing
Baseline
https://files.catbox.moe/o4r3jh.mp4
De-rope
https://files.catbox.moe/grs5xa.mp4
Slo-mo
https://files.catbox.moe/crn09s.mp4

Flux 3 prompt I was trying to mimic (though in its case I used a simple prompt https://files.catbox.moe/rf653a.mp4), even its API has this smearing issue that the MIT paper fixes.
>>
>>109502965
>>109502942
not for me bitches.

https://d.uguu.se/mlzEHovp.webm
>>
>>109502987
Actually, motion may be more funky because I took extra steps on the de-rope side
>>
I still think wan2.2 is overall better than h3. The main thing h3 has going for it is sound.
>>
>wan2gp added frame injection
ok, ok, now we are getting somewhere
>>
>>109503001
Patrick don't you have to be stupid somewhere else?
>>
>>109503001
You're crazy. H3 understands physics and object interactions that I never managed to get Wan to do.
>>
>>109502730
>straight out of diffusion.
https://files.catbox.moe/sddmtd.mp3
>Look in thy glass
>Shakespeare's 3rd sonnet.


>>109502742
>ace step 1.5 xl base, in case it's not obvious.
>This is the style prompt:
>jazz. angry screaming female singer.
>This is how I structured the lyrics prompt. idk, "virtual singer" doesn't seem to do much, but idk, I'm constantly throwing in random seasonings to see if anything is interesting. idk there's bass at the end, so maybe that part worked.
>>
>>109502993
yesterdays gens reheated. :vomit:
>>
sota local voice cloning and music gen when?
>>
>>109503008
>>109503010
go and try wan fun-VACE then come back.
>>
File: 1785446451639988.gif (259 KB, 270x200)
259 KB GIF
Local bros, we can do local on the cloud now
>https://videocardz.com/newz/modders-gain-full-windows-desktop-access-on-geforce-now
>free tier cant be exploited
>but paid versions of geforce now can
>hijack it and install lm studio
>5080h (RTX 5080) at their disposal
>>
>>109502978
oh anon, both the porsche and ferrari get GAPPED by a 4 door family sedan. It's time to realize EVChuds WON
>>
>>109503024
i'm never trying anything wan ever again you fucking retard.
>>
>>109502937
Here yoyu are
https://files.catbox.moe/44kp8a.mp4
>>
>>109503022
Literally the only things we're missing
I had great luck with omnivoice though
>>
>>109503048
Can cloud models even do sota voices yet?
>>
>>109503001
>>109503024
I have Wan fatigue. Every video model since Wan has used its architecture in some form. I'm never touching that shit again. H3 broke the cycle and is now my hero
>>
>>109502993
ZAMN! SHE'S 14?!

also please post your other ones, i missed them and noticed the reposts in the collage.
>>
>>109503022
We already have SOTA music gen (with LoRAs), and since Alibaba went closed source and likely won't release Qwen Music, we now wait for ACEStep v2 to answer that question for a music gen base model.
>>
>>109503061
>ban evading
>>
>>109502574
It's probably more of a 'damned if you do, damned if you don't' situation. I don't think that there will be a simple bubble burst, but more like a monkey paw one: a new GPU technology that is needed for a new type of AI, and while we can then probably buy the old GPUs for cheap, the new cool thing will be pricewalled behind the even newer and even more expensive new advanced GPU types and the ones who bought now will also be unable to use the new technology.
>>
>>109502367
Minimax has no gray safety box failures, and despite you setting up whole scenes, it's easier to prompt than Ideogram
>>
>>109503027
>local on the cloud
>>
>>109503027
cringe and poorpilled
>>
>>109503065
Other whats? Dancing Mays? I've got about 45, it's all basically a long troubleshooting process to figure out the best way to generate rhythmically synchronized (sexy) dancing.
>>
check this out, reference can copy the style of old ps1 games and add a character in the style.

https://files.catbox.moe/do0ecs.mp4
>>
thats ok, ill get bored of ai eventually and go back to playing video games instead so i dont need to worry about buying a super computer
>>
>>109503061
lmao besides the context of the vid yeah seedance 2.5 does look good
>>
>>109503092
yeah, those cool dancing may figures. share whatever you want. i just thought they were hot enough to hotglue tbqh.
>>
Seedance is rumored to be a 200b param model, so it HAS to look good lol
>>
you can use turbo lora as an upscale pass.
>>
>>109503061
Okay it looks good lmao
>>
>>109503097
yeah, it's kinda grim btw. nta, but a couple of days ago I got a hold of a seedance. And h3 is way worse than seedance in terms of video quality, camera movements, and... welol quite everything.
but h3 is local, and that's great.
>>
>>109503061
bruh.
>seedance cloud shit
Fucking rogue employees
>>
>>109503027
>>109503087
>not wanting to make dumb shit on custom to geforce now hardware
ngmi
>>
>>109503061
>1 hour to gen
>8 hours gooning like the pedo he is
.
>>
>>109503137
You mean 8 hours bypassing the safety filter. I sure wouldn't want to get caught with that
>>
File: q_xdg7zk.png (844 KB, 1024x1024)
844 KB PNG
>>
the end cutscene where she makes the "yuck" face and sound is Oscar-tier

>>109503113
>Seedance is rumored to be a 200b param model
look what they need to do to mimic 1.5x our power

>>109503120
>Fucking rogue employees
actually like 60k of stolen keys iirc
that vid probably cost at least $1000 on API to make

>>109503119
>h3 is way worse than seedance
but the thing is, i don't NEED seedance 2.5 to make a girl posing, or basic gooner shit. so there is actually a "good enough" threshold for a significant chunk of my uses with AI and I think H3 crosses that, and there's no going back from that

>>109503137
kek, but seedance clips at 720p take around 3-4 minutes so i believe it. idk what the rate limits were fior him since the website he used thats watermarked in the corner switched to whitelist-only so maybe he was doing 1 video at a time, maybe 4 videos at a time idk
>>
>>109502604


I keep hearing this and I keep trying to insert a single mom I know into deviant shit and can't get it to work to save my life.
>>
is he right or is he right?

https://files.catbox.moe/ie449b.mp4
>>
>>109503103
Alright here's a randomly curated selection I guess. You'll be sick of Toxic like I am after watching them lol.

https://n.uguu.se/luFBwOUd.webm
https://d.uguu.se/wXcBLqqH.webm
https://h.uguu.se/GuryEQBp.webm
https://n.uguu.se/vGKpqGwC.webm
https://n.uguu.se/oqghfnYa.webm
https://d.uguu.se/DcTnvsxi.webm
https://h.uguu.se/WPRWMazd.webm
https://d.uguu.se/hBCKiFGA.webm
https://n.uguu.se/AKPveQhW.webm
https://d.uguu.se/RhgwbbBN.webm
https://h.uguu.se/qieBhMBN.webm
https://n.uguu.se/iKnSLvly.webm
>>
>>109503061
>>109503097
>>109503113
>we're gonna have shit like that locally in 2 years
future's looking bright, maybe less so for my wallet with how expensive upgrades are gonna be, but trust the plan.
also lol'd hard when she saw his dih that was actually pretty well made despite the content (which i FULLY disavow for any glowies reading this)
>>
>>109502220
>rtx 6000 pro
post workflow please
>>
>>109503168
nothing shitdance can do, can compare to what the reference minimax model can do. you can do LITERALLY anything, which is something a rigid, censored, static model like shitdance can't do.
>>
>>109503027
is this some sort of poorfag joke i don't understand?
>>
File: MiniMax_H3_00126.mp4 (2.54 MB, 960x720)
2.54 MB
2.54 MB MP4
>>
>>109503168
>I think H3 crosses that
fair enough by the way. the only things I make is gooner shit. And it's better to use local models for that. Hope that 2k upscaler they have will make wonders.
>>
>>109503169
>insert official ref model prompting guide into claude
>generate skill .md file for a prompt enhancer
>feed skill to gemma 4 ablit
>tell it exactly what you want it to do, and exactly what inputs you have to work with
>check prompt afterwards and adjusting timing/content to taste
>>
>>109503061
>genning this off local
Absolute balls. I'm scared even using grok to fix my prompts.
>>
>>109503193
Can I use skills in llama-cpp web UI?
>>
>>109503174
kino, thanks.
>>
>>109503200
A skill is just a system prompt. Use whatever. I use LM Studio myself.
>>
For anyone using ComfyUI, when I divide a line into two, is there a way to use switch to decide which line I want to go?
>>
File: 1769446700294860.mp4 (3.7 MB, 544x960)
3.7 MB
3.7 MB MP4
its a bumpy ride
>>
>>109503174
fuck off spammer
>>
might switch to int8 vae, VAE encode is FUCKING SLOW
>>
>>109503174
Why are you obsessed with pokemon girls?
>>
autism alert:
>>109503215
>>109503211
>>
Sulphur is at $9855/$10,000. That was fast.
>>
>>109503119
Yes because you are a Vramlet shizo that use mini max with 0.4 turbo slop lora and eassy cache. In 2mp the model can easily reach that level of detail and you know, the model can do the part you corpo cannot.
>>
>>109503213
but it makes outputs look like shit. chill man, it doesn't even make it faster.
>>
>>109503224
never underestimate coomers
>>
Is the Pokemon autist and the Rocketgirl fag the same person?
I was hoping he finally realized no one outside of /vp/ cares about that shit at this point
>>
>>109503197
>Absolute balls. I'm scared even using grok to fix my prompts.
it was with stolen keys. do not try this at home.
>I'm scared even using grok to fix my prompts.
look into zero data retention openrouter endpoints or ask AI to explain them to you. these are what companies use for provable compliance for processing healthcare records and stuff so its actually private. if you trust openrouter's providers that say they don't keep logs at all you can just use openrouter and any model you want and pay with crypto
>>
>>109503225
you forgot that minimax h3 also has an api version
>>
File: 4557845437894.png (18 KB, 900x806)
18 KB PNG
where is kino?
>>
>>109503229
It makes the decoding faster, but yeah, it also lowers the quality
>>
>>109503232
Or just run things locally so you don't have to trust a single third party
>>
File: 1759224735455493.mp4 (3.43 MB, 1216x672)
3.43 MB
3.43 MB MP4
>>109503241
>where is kino?
im out of ideas and getting tired of first person pov for sfw stuff

>>109503251
i was gonna suggest that too like qwen 3.6 or something but you're probably running h3 locally already
>>
>>109503256
>im out of ideas and getting tired of first person pov for sfw stuff
mind sharing the prompt? i wanna start doing foot soldier kinos
>>
>>109503225
>In 2mp the model can easily reach that level of detail and you know
How much VRAM would H3 need for like a 15 sec video at 2mp?
>>
I'm jealous of you anons, I'm so autistic I'm updating everything properly, and checking every custom node for security issues before even trying h3.
>>
>>109503268
yes
>>
>>109502604
I used grok to gen every member of blackpink but with fat asses then superimposed an image of myself in them and have been going hog wild with i2v. Seriously feels like you're a god. The shit I've made them do is diabolical.
>>
File: chad horse.jpg (35 KB, 736x971)
35 KB JPG
>>109503270
>checking every custom node for security issues
why do you need to do that? what are you up to my man?
>>
File: q_a22ji5.png (1.13 MB, 1024x1024)
1.13 MB PNG
>>
This is 100% what seedance 2 does
https://matlowai.github.io/ComfyUI-MAINodes/#featherweight
Fixes the blurryness completely
>>
>>109503270
Anything to report yet?
>>
>>109503270
i use someone elses ui so all i only have a single git history that i review before i update
>>
>>109502333
why can i not have a asuka or hermoine gf can someone pls explain why does the universe not allow me this?
>>
File: what did it cost.png (28 KB, 1024x316)
28 KB PNG
>>109503279
im gonna cry
>>
>>109503279
someone translate this Claude speak into actual human words
>>
>>109503241
is in /sdg/
>>
With R2V using an audio reference for voice, do any of you have the issue where an undesired noise/vocalization plays at the start of the video?
>>
debo this is no time for jokes
>>
>>109503297
you wish
>>
File: output_small.mp4 (3.89 MB, 2048x1142)
3.89 MB
3.89 MB MP4
Had no idea 4chan now takes mp4
>>
File: H3_00010.png (1.27 MB, 1216x672)
1.27 MB PNG
>>109503266
>mind sharing the prompt?
ask your AI for live action gopro stuff e.g.

Live-action, cinematic, first-person POV, a wide GoPro-style lens frames a windswept riverside meadow where armored barons in chainmail stand in a grim ring around a campaign tent, banners snapping overhead.

>>109503270
gotta offload your autism to an AI
>>
>>109503270
>>109503280
If someone put some secret backdoor into a node that phones home I think it would have blown up already. I fully expect glowies have some undetectable shit already embedded into every possible base local install.
>>
>>109503113
It doesn't even look that much better though, the shills will just be shills

https://xcancel.com/magic_ai_skill/status/2084437127618826257#m
>>
>>109503313
thx
>>
>>109503268
more than 40
>>
>>109503321
And that's the thing. Twitter etc... are filled with Seedance shills, you can tell right away because they do not prompt engineer the model they're comparing against and claim it automatically loses. Flux 3 is better than Seedance 2.5 and is far less slopped, there's no doubt about it.
>>
which comfy "get image size" node actually shows the numbers in the node, so you dont have to connect more shit to it?
>>
>>109503313
also if you need ideas, do first person view of big disasters like volcanos erupting or something like the hindenburg crashing down
>>
>>109503335
Yep, no contest, F3 is the video model and it's likely magnitude times smaller than Seedance
https://xcancel.com/isocialwebseo/status/2085820290437697803#m
>>
>>109503350
>video model
best* video model
>>
>>109503350
flux 3 dragon looks like how to train your dragon dragon
>>
>>109501448
that's how I became a rebetikochad
https://www.youtube.com/watch?v=WpDZ1uVDbt4
>>
People seem to be completely forgetting these:
1 - People are running H3 local with shit the severely lowers the quality of the outputs, 0.4MP, those "speed optimization" nodes, quants etc
2 - They haven't released the 2K upscaler yet
3 - That movie-like quality Seedance has may be achievable with a competent fine-tune. No has has done a large scale fine-tune on H3 yet. Training on real high-quality movies may add the "grainy" texture, make things sharper and reduce the "plastic" look. H3 out of the box is pretty slopped, but the underlying model is good.
>>
>>109503384
this exact same post was made about wan 2.2 when it first came out
>>
>>109503384
Oh, and training on movies can improve the facial expressions too
>>
>>109503384
I make high quality works see
>>109503308
I only gen at 1mp and upscale with rtx it's a cope until they release the upscale but we have flexibility. The turbo loras are also fine, we're like week 1 and look at all the crazy shit we have
>>
>>109503402
Wan2.2 was and still is really good. The problem is that the poorfaggotry made it look worse than it actually was.
If Seedance 2.5 got open-sourced, most people's gens would look far worse than the API as well
>>
>>109502604
How good does this work exactly? I tried it a bit but it just kind of cut to a slightly altered version of the same video. Even with generated prompts. Can you make so it just copies the motion loosely and apply it to an entirely new location/angle/character?
>>
>>109503418
did you define the motion as its own subject? then you say its attribute_transferred to another subject
>>
>>109503412
>If Seedance 2.5 got open-sourced, most people's gens would look far worse than the API as well
it wouldnt because poorfags wouldn't bother with it because H3 exists
>>
this is legit amazing. I have not tried continuing video before
https://www.reddit.com/r/StableDiffusion/comments/1viwvuq/better_avoid_saul_2_minimax_h3/
>>
>seedance
https://files.catbox.moe/1mo32o.mp4
>>
>>109503430
okay i still want you to kill yourself for continuing to link reddit, but credit were its due this is fantastic.
>>
>>109503442
saaar we all use the reddit here and we LOVE it
>>
>>109503436
ah yes the classic youtube gamer, "the barely perturbed gaming geek".
>>
>>109503215
they're very fuckable
>>
>>109503224
My only issue is if it will look good or if every donor added the same "professional porn with botoxed milf" shit.
>>
Which of the speed-up options for h3 won't kill my gen quality? I'm just running sage attention 2.2 currently.
>>
>>109503458
I gave a bunch of real amature stuff
>>
>>109503430
I'm guessing those were genned at maximum resolution with no sage or any type of speed cope cause they're pretty good
>>
>>109503275
Yeah, my personal files, secret in env and so on.

>>109503280
So far so good.

>>109503313
>gotta offload your autism to an AI
I use gpt pro model, pretty amazing model, but I verify everything it writes, and it takes time.
>>
>>109503461
i think none of them are bad on their own, but if you stack all of them up then it starts to get worse
>>
>>109503461
>>
Haven't installed H3 yet. Do weights actually NEED to be in vram, or can I offload to DDR5 for longer gens?
>>
uhhhh...animator status?
this is just with the turbo lora too
https://files.catbox.moe/uhlb1k.mp4
>>
>>109503308
crazy gen
>>
>>109503465
Nice.
>>
>>109503461
Use the pruned model, and outside of sage 2, literally nothing.
>>
>>109503474
>voice actors still safe
>>
>>109503308
man, h3 is magic
looks indistinguishable from some old anime
AI porn slop spam is about to have a huge quality increase
>>
>>109503472
Should I not be running the H3 Mem Eff node?
>>
>>109503492
well, voice actors that don't have degenerative brain diseases anyway.
>>
>>109503503
depends on your vram but the normal one is better objectively
>>
my brain feels funny ever since I started genning with minimax.
>>
>>109503514
I thought he was just Welsh
>>
>>109502367
You can't even draw a piece of bread with ideogram with a simple prompt. If doing something even so mundane and simple requires esoteric bbox fuckery, it's not worth the effort.
>>
>>109503519
weird innit
being able to translate both your brain's thoughts and dickvision into a 1:1 video anyone else can see
>>
Had to wage slave and took 2 days break. So what's the meta for quality/speed balance? For me its KJ Sage patch. All other cope caches seems to just make faces and hands blurry during dancing.
>>
I want to use the turbo lora as an upscale pass, but fuck VAE encode is like 2min at .8mp
>>
>>109503471
>>109503472
>>109503490
Thanks anons.
>>
File: file.png (359 KB, 3350x922)
359 KB PNG
>>109503534
If you're on a 50 series
Quality degradation, especially on higher mp, is almost non-existent
>>
https://litter.catbox.moe/bxk84q.webm
>>
>>109503541
Thanks, I'll load up some gens and report back.
>>
>>109503541
>sol AND spectrum
post results, i find it hard to believe stacking the two doesn't result in degradation.
>>
https://www.reddit.com/r/StableDiffusion/comments/1vj79sz/minimax_h3_some_test_spectrum_lightx2v_larryvrh/
optimization comparisons. It looks like larryvrh Ema v4 600 LoRA is the best over all speed vs quality. Spectrum is slightly better quality but is slower
>>
>>109503551
Sure, let me gen something sfw first
>>
We also can tell if you're frauding anon, me and >>109503551 will know
>>
>>109503541
i still have no idea if sage attention 2 works on amd or not
>>
>>109503554
>heretic text encoders
FUCKING DIE
>>
so h3 is what people who can visualize things in their head feel all the time? literally imagine a scenario and see it like some kind of magical performance
damn I'm glad I was born to see what it looked like
>>
>>109503060
>sota
what's that?
>>
File: not my screenshot.jpg (383 KB, 2159x880)
383 KB JPG
>>109503535
>apply film grain
>no need to cope with uprez/detail fixing
>>
>>109503546
I am inversely disappointed by pokemon games vs pokegirls' designs. They're always spot on.
>>
File: MiMx-IMG_00111.mp4 (3.58 MB, 832x1120)
3.58 MB
3.58 MB MP4
>>109503519
2 minutes for 5seconds feels too long.
>>
>>109503554
Using ablit TE reduces the quality of your outputs. It also doesn't "uncensor" them. All it does is give you worse results than you would have had with the proper TE.
>>
>>109503558
i await with (master)bated breath.

>>109503575
wow he's literally me
>>
File: output_small2.mp4 (3.96 MB, 2048x1142)
3.96 MB
3.96 MB MP4
>>109503554
Is that plug and play if not I'm sticking to the one KJ put up.
>>
>>109503571
You are missing separating the layers into r, g, b and adding subtle noise to each layer and then combining them back together
>>
>>109503546
FUCK OFF FUCK OFF
>>
>>109503570
state of the art
>>
>>109503586
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/blob/main/minimax_h3_turbo_v4_step600_ema.safetensors

Its just this. Its way better than the light lora / spectrum
>>
>>109503569
Yes but my visualization is like an old tube TV. Now we can generate things in high def.
>>
Does Spectrum have an effect on sound quality? Specifically dialogue.
>>
>>109503571
turbo lora rapes coherence and prompt following.
this way I can gen at low res with zero cope nodes to get maximum model performance. than upscale to get rid of all the ugly artifacts.
>>
>>109503597
I'm sure there are people who can imagine in 4k or something
>>
>>109503594
idk this shit doesn't work for me at all.
is it just bad for ref2video or something? Can't get it going
>>
>>109503601
not this one
>>109503594
>>
FRESH
>>109503609
>>109503609
>>109503609
>>109503609
>>
>>109503599
Used to for the ref model, but the new version fixed it or heavily mitigated it. Sounds good now. Seems it broke the preview though, on both the sampler and KJ's custom preview for Minimax. Least it's not working for me and a few other people looks like, least on ref. Haven't tested base yet
>>
>>109503575
I don't think that's what "eating ass" means.
>>
One sec lads, I'll bake the non-debo thread.
>>
>>109503608
Yes this one >>109502449
>>
>>109503606
it needs this loader
https://github.com/Larryvrh/ComfyUI-MiniMax-H3-Turbo
>>
>>109503611
>trollbake
again?
>>
>>109503594
I don't want to install a custom node, or is he just providing one for people that have no idea what to do?
>>109503611
autistic miserable wretch
>>
>>109503624
mrbones.mp4
>>
>>109503615
make a fucking collage. if you can't I'll make one.
>>
>>109503626
you need it atm, pruning loras does not work well
>>
>>109503629
You have 2 minutes, autist.
>>
>>109503611
get a job sharty, stop wasting your time with that inane shit
>>
>>109503631
I'll stick to the kj one, I don't really mess with nodes not made by a few trusted sources
>>
>>109503637
ok its coming
>>
>>109503640
you usually come that fast?
>>
>>109503643
no, but the guy he's blowing does
>>
inpainting ace step 1.5 xl base. Loads of fun.

Doing Sonnet II of Shakespeare.

I reckon I'll do them all, if I live long enough to finish it out.
>>
>>109503637
>>
>>109503646
lol lmao
>>
im learning. what do you think of my first gen for h3?
https://files.catbox.moe/rdfbyu.webm
>>
stfu faggots you do not disrespect bakers unless he adds his own gens to it
>>
WHOEVER IS VIBECODING SPECTRUM, FIX THE FUCKING BROKEN PREVIEW
>>
>>109503639
https://litter.catbox.moe/j6ge0eohu0r678hj.mp4
>>
its up
>>109503609
>>109503609
>>
>>109503651
way past cool bro
>>
>>109503210
based vanilla coomer
>>
>>109503649
was #8 posted here? holy fuc
>>
>>109503648
reminder this week's gens
india diss:
https://files.catbox.moe/dioqb9.mp3
sonnet III
https://files.catbox.moe/sddmtd.mp3
sonnet I
https://files.catbox.moe/w2ysvt.mp3
>>
>>109503210
>>109503663
nvm found it
godtier
>>
>>109503659
gotta go fast
>>
Move
>>109503671
>>109503671
>>109503671
>>
>>109503649
>give a priest the job of driving one nail
lmao accurate
>>
>>109503660
Remember when all it took was a nice pair of boobs?
>>
>>109503676
Yeah, if a single lady with huge boobs and of modest face says hi to me at church tomorrow, I'll punch her. Disgusting whore!
>>
>>109503676
just be a teen again
>>
>>109503656
>when you look into the workflow and there is more spaghetti than ten ai-generated Will Smiths could eat
>>
>>109503541
spectrum is supposed to be at the end of the chain, dummy
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3#workflow-placement
>>
>>109503907
Thanks, I'll try it
>>
File: 1785963412471182.png (135 KB, 474x532)
135 KB PNG
>>109503430
>this is legit amazing
>>
File: 1762810729868472.png (1.12 MB, 1254x1254)
1.12 MB PNG
>>109504222



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.