[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


Discussion and Development of Local Image, Video, and Music Models

Previous: >>109507737

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
>mfw Resource news

08/09/2026

>Kroma v0.2 — Krea 2 fine-tune (full model)
https://huggingface.co/lodestones/Kroma

>krea2-turbo-bbox
https://huggingface.co/jimmycarter/krea2-turbo-bbox

>Kroma v0.2 Quant
https://huggingface.co/silveroxides/Kroma-Quant/tree/main

>Spectrum for Ideogram 4
https://github.com/Nif00/ComfyUI-Spectrum-Ideogram4

>ClipProj — MiniMax H3 conditioning from a Qwen3-VL-4B
https://huggingface.co/NicoLab28/ClipProj-MiniMax-H3

>ComfyUI-SigmaSync-LoRA: Sigma-aware model-only LoRA strength scheduling
https://github.com/capitan01R/ComfyUI-SigmaSync-LoRA

>NexusBTA v0.2.44 adds MiniMax H3 support
https://github.com/JpAndreBTA/Nexus-BTA/releases/tag/v0.2.44

>Experimental MiniMax H3 single-image VAE
https://huggingface.co/Mamad8/MiniMax-H3-Image-VAE

>MiniMax H3 REF2VA w4a8
https://huggingface.co/realrebelai/Rebels_w4a8s

08/08/2026

>Kijai: MiniMax H3 Ref Lora Rank 256 bf16
https://huggingface.co/Kijai/MiniMax-H3-experimental/tree/main/loras

>MiniMax H3 at native fp16 on pre-bf16 GPUs (V100 / Volta)
https://github.com/Amduraznak/minimax-h3-fp16-fix

>Cosmos3-Nano-WebUI: Self-hostable API + Web UI for Cosmos3-Nano quantized fp8 and nvfp4 checopoints
https://github.com/fengwang/Cosmos3-Nano-WebUI

>R9700 AI Pro — ComfyUI / MiniMax-H3 speed patches
https://github.com/charlie12345/R9700AIProComfyUIPatch

>MiniMax-H3-Pruned-GGUF
https://huggingface.co/Abiray/MiniMax-H3-Pruned-GGUF

08/07/2026

>OpenLayer v0.13.0-alpha — ComfyUI in Photoshop, free and entirely local
https://github.com/MehranMarxian/OpenLayer/releases/tag/v0.13.0-alpha

>LIGHTX2V 4-step Turbo Minimax H3 lora
https://huggingface.co/lightx2v/Minimax-h3-Turbo

>LIGHTX2V MiniMax-H3 T2VA Prompt Rewriter LoRA
https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA

>Sage Ready: Local-only installer and readiness checker for SageAttention
https://github.com/CosmicFungi/Sage-Ready

>Wan 2.2 Animate 2 14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B
>>
>mfw Research news

08/09/2026

>MDTD-ArtIR: Benchmarking Image Editing and Restoration Models for Art Image Restoration under Texture-Overlay Degradations
https://arxiv.org/abs/2608.00736

>Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
https://arxiv.org/abs/2608.03284

>Test-Time Curriculum for Open-Set AIGC Detection
https://arxiv.org/abs/2608.00559

>Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
https://arxiv.org/abs/2608.00094

>Visual Anchoring in Diffusion: Multimodal Zero-Shot Skeleton Action Recognition
https://arxiv.org/abs/2608.04623

>Free-Lunch Augmentation by Revisiting Diffusion-Based Data Generation for Cross-Domain Few-Shot Object Detection
https://arxiv.org/abs/2608.04394

>Controllable Clothing: Precise Labels and Generation for Virtual Try-On with Latent Diffusion Models
https://arxiv.org/abs/2608.05834

>EulerLoRA: Rank-Driven Jump Dynamics for Calibrated Parameter-Efficient Fine-Tuning
https://arxiv.org/abs/2608.01142

>DeVIT: Low-Power Vision Transformer Acceleration Using Delta Computation
https://arxiv.org/abs/2608.01343

>Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning
https://arxiv.org/abs/2608.00994

>MiniWorld: Democratizing the Training of Video World Models from Scratch
https://arxiv.org/abs/2608.01127

>GVCCTurbo: Rate-Compute Quality Scheduling for Codebook Driven Generative Compression
https://arxiv.org/abs/2608.03517

>Hi-Token: Hierarchical Coordinate Tokenization for Generative Visual Grounding
https://arxiv.org/abs/2608.03471

>Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression
https://arxiv.org/abs/2608.02134

>OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
https://arxiv.org/abs/2608.03812
>>
im in the middle of baking the collage and then some retard always snipes it. whatever
>>
File: patientgenner.png (47 KB, 962x454)
47 KB PNG
correcting my slight early morning groggy retardation here >>109509081
it was 8sec, not 10. 10 ended up putting me at 59s/it from 43s/it. I imagine getting my extra 32 gigs of ram could help the offloading and put me back in the forties.
Still a fine speed IMO i don't mind patiencemaxxing, the quality's good enough.

>>109509231
5060 ti 16gb

>>109509254
reduces VRAM peaks, so depends on your situation/how much vram you use and if you're even paging out of system ram. which i am either way, 32 gigs is not enough.
>>
>>109509313
>>109509318
WARNING MALWARE LINKS! Take care anons!
>>
>>109509348
Catbox workflow?
>>
File: screenshot.1786294004.jpg (851 KB, 1777x889)
851 KB JPG
qwen3.6 uncensored works great for h3 prompt structuring. just make sure to disable the thinking model
>>
File: 174046CUI_00001_.png (1.11 MB, 1216x832)
1.11 MB PNG
>>
>>109509343
seriously, I get baking to not get trolled, but I'm really getting tired of not having a collage.
>>
man finds out Flux 3 cant generate a good Miku

https://files.catbox.moe/ls9sj4.mp4
>>
>>109509343
You always pretend like you're shocked when a new thread is made after the previous hits its bump limit. You can always see the reply count. Just start baking your collage earlier if it's this important.
>>
Did any of you guys save that skateboarder girl shortfilm from yesterday? I didn't save the litterbox before it expired. Honestly blew my mind what you can make with AI now.
>>
>>109508984
This is worse right?
https://h.uguu.se/VlWFUWhb.mp4
>>
>>109509313
>>109509318
Go back to your bot ridden containment general loser
>>
>>109509396
early baking is bad, the only reason we do it is because we'll get a troll bake otherwise.

But I think we can stop doing it, seems like the jannies are deleting the troll bakes now.
>>
>>109509400
hate uguu and litterbox like you wouldn't believe
>>
supplying last frame instead of first frame is also very fun for H3 img2vid
>>
>>109509400
nice try glowie.
>>
File: t2v_MiniMax_H3_00009_.mp4 (1.52 MB, 672x1216)
1.52 MB
1.52 MB MP4
>>
Tell me about debo: why does he keep trying?
>>
File: Barclay holodeck.png (827 KB, 1024x1024)
827 KB PNG
>>109509301
>I have asked this 3 times now, once in a language thread with similar general name and once in a troll thread.

I ask again

What is the best image editor on ComfyUI? Is it still Qwen? My PC can’t handle it. I wonder if a less VRAM-intensive version is out. My PC setup is 12 GB VRAM and 16 GB RAM.
>>
>>109509423
Honestly I get why they can't host ridiculous amounts of data forever, I just wish it would default to at least like 1 month of retention, or maybe a week of inactivity or something.
>>
What sampler have you guys settled on in Minimax? res_multistep?
>>
>>109509441
res_multistep 20 steps is good enough for a start, if your motion is blurry or prompt isnt followed too great then increase to 30-50
>>
>>109509433
Flux.2 Klein. Try the default Comfy workflows for it.
>>
>>109509433
I thought Klein had displaced Qwen for editing. (Get the kv version, it's supposed to be faster by avoiding redundant recomputes.)
>>
>>109509433
>COMPUTER!
>CREATE A SIMULATION WHERE NIVIDIA, OUR RULER, FAILED TO CLAIM 95% OF THE MARKET
>SHOW ME A WORLD WHERE AMD AND INTEL SURPASSED EXPECTATIONS AND BECAME THE DOMINANT COMPANIES
>"I'm sorry am an ethical llm model and cannot fulfill this request"
>FASCINATING
>>
>>109509457
>>109509455
Thank you for your help , may you get perfect generations.
>>
>>109509408
https://h.uguu.se/VMFgsiYa.mp4
This one looks better.
>>
>>109509369
>just make sure to disable the thinking model
why? so it's quicker or does the thinking do something bad / too much context waste?
>>
user discovers flux 3 is inferior to minimax

https://files.catbox.moe/us0uxu.mp4
>>
>>109509511
kek
>>
Hey bros, these threads move so fast the archive is hard to use. What's the current poorfag H3 cope meta for vae, TE, and the main diffusion model? I have a 4000 series card.
>>
>>109509418
It's all troll bakes as long as the drama trannies like you throw a fit in every thread
>>
https://n.uguu.se/ArXCeUsi.mp4
>>
>>109509517
threads are hard to use because of the aforementioned drama trannies. just use reddit to find speedups for now since that's all these retards are parroting anyways
>>
how do you get no dialogue? sometimes my gens be speaking straight gobbledygook
>>
Do you guys think WAI-Illustrious is obsolete? I still use it for mask detailing (never found a good workflow for anima), and it has some loras I still like.
>>
File: MiniMax_H3_00761.webm (1.07 MB, 1184x896)
1.07 MB
1.07 MB WEBM
>>109509408
>>109509484
Yeah, the second one is better. It captured the cel jitter
>>
>>109509542
So good
>>
>>109509530
Lol, I'm the retard, it took me all weekend to set up sage 2.2.0 and sol attn. Ubuntu 26.04 based distros do not fucking like sage.
>>
File: screenshot.1786299499.jpg (104 KB, 816x311)
104 KB JPG
>>109509505
thinking mode wastes more time and doesn't improve the output enough to be worth it. cost analysis
>>
>>109509542
This is just turbo 8steps
https://n.uguu.se/wvAEBOpT.mp4

This is my problem with turbo, it does very sharp, but coherence goes out the window.
>>
>>109509534
if there is some dialogue that you want, use the format
>X says <d>[English] your text</d>
instead of
>X says "your text"
This will fix it.
>>
File: instaboner.jpg (307 KB, 1389x1392)
307 KB JPG
fffuuuckk muh dih
their image gen model will be earthshattering - if they don't rugpull us anyway.
>>
File: 1785705869050094.gif (937 KB, 800x450)
937 KB GIF
>>109509489
theres also the fact that intel got a money injection from nvidia as with the american government owning 10% so its not surprising nobody knows what the fuck is going on with intel and their gpu division but i like that we have a third option at all
support as you see fit, fuck that guy saying buy stocks, stocks is just gambling like ai is, but supporting the software side of it will deal with some of the other pains long term
>>
File: screenshot.1786299998.jpg (70 KB, 534x581)
70 KB JPG
Character Reference Loader node ready! along with forked MiniMax H3 Reference to Video node that takes an image list and matches them to the ref_images
>>
>>109509609
give her fatter tits
>>
check this out
>spectrum off
>600ema turbo lora at 0.66
>12 steps
>euler/beta
>>
>>109509570
Why are you larping like you are the one who made that
>>109509609
Do you know if it matters if your references are higher resolution compared to the output dimensions? Does your node account for that, or is there a different way of doing it without losing out on quality?
>>
https://github.com/NikoDemon80/ComfyUI-H3-Motion-Context

can chain clips with this
>>
>>109509564
plugged a wrong spaghetti so this one is 0.5mp both passes.
https://d.uguu.se/elWRtJNv.mp4
>>
>>109509631
slop
>>
>>109509612
absolutely not

>>109509619
>Do you know if it matters if your references are higher resolution compared to the output dimensions?
from my experience, yes. i don't have solid evidence to prove it though, just my own anecdotal notes. make sure ref_image_size is set to max. this allows the model to get small key details it would otherwise miss such as accessories. gen times will increase a bit but it's worth it, after all, the point of ref2va is the reference.
>>
Been testing this:
https://huggingface.co/drbaph/MiniMax-H3-Turbo-Lora-ComfyUI/blob/main/minimax_h3_turbo_v4_step600_ema_pruned_comfyui.safetensors
At Strength 1.0 + Euler + beta + 8 steps
Feels much less erratic than v1. Still there is some blur and sometimes it adds weird details to video but feels much closer to usable.
Anyone else tested it? Anyway to improve its outputs?
>>
The new spectrum node upgrade broke the preview progress of sampler.... Is ogre
>>
ok who's the anon who said to use the int8 q8a8 model loader node, still testing it but as far as I can tell it's a huge speed boost for no discernible quality loss

may you always get trips anon
>>
File: 11317754990074.mp4 (3.51 MB, 736x576)
3.51 MB
3.51 MB MP4
>>>/wsg/>>6211018
>>
>>109509674
update it again
>>
File: 1763368591717204.png (62 KB, 1010x582)
62 KB PNG
>>109509679
i tried it out as well and i'm genning 1mp now. amazing
>>
>>109509686
ah shit >>>/wsg/6211018
>>
JC Denton runs into someone familiar in the city

https://files.catbox.moe/5d7udj.mp4
>>
>>109509679
>>109509698
official nodes support it if ur on cu130
>>
>>109509698
>model type flux2
is that right?
>>
>>109509570
Content aside. How did you make 2 minutes animation like this ?
>>
>>109509711
no idea, it just works
>>
I need to add more swap. VAE memory management still fucked.
>>
>>109509716
By genning four 30 second clips and editing them together?
>>
>>109509696
I did it still broken
>>
>try Minimax
>30 mins for a 4-second 0.2 MP clip on my iGPU
Not practical, but possible! I thought it was doing 20 steps for a single frame at first, so I'm glad I let it finish rather than abort.
>>
File: 1771387245473543.png (471 KB, 736x576)
471 KB PNG
>they trained minimax on Seinfeld but not Frasier
>>
>>109509724
>>109509732
What app for combining videos ? Might do it someday when Minimax matures
>>
>>109509609
Post links to the images please. I need Kanna butt...
>>
>>109509708
https://files.catbox.moe/nv7xmf.mp4
>>
>>109509751
https://x.com/ai_daihuku23/media
>>
>>109509740
>What app for combining videos ?
anistudio
>>
>>109509740
Brother if you can't figure out how to stich 4 clips together you ain't making all that, sorry.
>>
You think it's possible to actually make full length episodes or even a film?

>>109509758
Thanks anon. Kanna sex.
>>
File: 447.jpg (29 KB, 394x474)
29 KB JPG
>What app for combining videos ?
>>
>rando anon asks retarded question
>my brain; awesome another reason to dogpile

>>109509740
dumb fucking retard lol
>>
>>109509739
nobody wants to gen that niggas big ass forehead
>>
>>109509739
indeed. Friends is missing a couple of characters as well.

Also, two things. I posted it on a previous bake, but I still want some answers. Does anyone else get a few bad gens (bad as in not following the prompt correctly) and then, after it gets it right once, the next gens are generally correct? Don't know what may be causing this.
Also: how in the fuck does h3 nail character's facial movement when the model was not specifically trained on a certain character? For example, I genned a few of the kike lord just starting from a still picture and it pretty much nails it. I know for a fact the model is not trained on his face, so how in the fuck can the model do this?

https://files.catbox.moe/14dbtv.mp4
>>
ayy yoo cuh what app to play on my iphone
>>
>>109509778
If I had the autism and didn't have a job I would absolutely spend my week doing this

I might anyway
>>
>>109509778
nta but I wonder if you could extend a shot by using the last frame and use img2vid. might be able to stitch some shots together for a seamless long shot
>>
>>109509761
>>109509765
>>109509772
Using adobe after effects is overkill. I asking for something simpler for AIslop
>>
File: 1775815630459655.mp4 (1.78 MB, 576x928)
1.78 MB
1.78 MB MP4
>>
>>109509792
>after effects
lmao, brown nigga what are you even on about just use davinci resolve you el tardo its FREE
>>
File: 248872065185582.mp4 (3.84 MB, 736x576)
3.84 MB
3.84 MB MP4
>>109509686
Adding a reference video helps with the dance moves, but bleeds in the style. Maybe it could work with a more specific prompt.
>>>/wsg/6211025
>>
genning at 0.4 mp really hurts the quality.
I need to gen at 0.6 mp minimum but it takes forever to gen
>>
>>109509809
img2img the reference to match the style
>>
>>109509739
>uncshows nobody nows
>>
>>109509804
Still too overkill. I want something simpler similar to "Webm for Bakas" i use
>>
>>109509818
just go visit the retard store and buy the retard editor
>>
being able to get 15 seconds of content in under three minutes customized with any subject imaginable for my very specific fetish is absolutely wild
>>
>>109509821
okay where are the retard store and the retard editor ?
>>
File: nog.jpg (32 KB, 750x559)
32 KB JPG
>>109509818
>Webm for bakas
>>
So far I'm not seeing a difference between the text and reference model, it seems to also be able to handle references well when properly prompted including style transfer.
If not for the loras I would stick to the ref model but this isn't adding up
Please advise
>>
>>109509824
low resolution fetish?
>>
File: 1764858407520583.mp4 (709 KB, 576x896)
709 KB
709 KB MP4
>>109509802
Based. Prompt?
>>
File: 1762317310902936.jpg (94 KB, 1136x521)
94 KB JPG
>>109509832
Great tool. It just works
>>
>>109509698
Doesn't work with lora(s?). tried with turbo and getting static
>>
Lets just admit. We all suck at this and this guide https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md is retarded.
>>
>>109509851
lora mode needs to be set to stochastic
>>
>>109509842
It's just a wrapper for ffmpeg nigga, an outdated wrapper for an ancient version of ffmpeg
>>
>>109509802
>>109509841
my down syndrome retarded low IQ queen returns
pls catbox these gens
>>
>>109509860
I want a GUI. not fiddling around with cmd
>>
>>109509882
You can have chatgpt slop you up something modern that'll do the exact same thing in about 30 seconds
>>
>>109509840
I admit to not caring about visual fidelity as much as the average viewer but .3mp and the firstblock cache get me to a very solid place. Taking firstblock off it's still under five minutes. I get that people can get faster gens but this is good enough to blow my mind still
>>
>109509882
Why are you even here if you can't ask google to give you a one shot command
>windows
>AI
keeeeEeek
>>
>>109509882
tell ai what you want so it makes you a .bat that you can double click nigge
>>
>>109509898
>>109509895
sigh.... you guys really rely on AI on literally everything
>>
>>109509905
have you been living under a rock? modern AI is better than 90% of programmers.
>>
>>109509905
Bored on a sunday?
>>
File: 1759649029404476.png (10 KB, 78x83)
10 KB PNG
https://civitai.red/models/2839513/male-ass-h3

that preview image made me laugh my ass off (nsfw)
>>
>>109509911
>>109509916
Im not going to pay Sam Altman to create a simple app
>>
>>109509905
thats fatherless behavior
>>
need more prompt donations to test my 2pass WF.
>>
>>109509925
Are you too stupid to run the model locally?
I think /sdg/ is more up your speed, it's a place for low ability posters
>>
File: 1761136173922691.png (2.77 MB, 1248x1824)
2.77 MB PNG
>>
>>109509925
Fine anon. I'll do it for you. Tell me what you want. Just combining videos in a GUI? That's it?
>>
>>109509936
>Telling people stupid without providing a solution
Lol
>>
Best text model for interpreting the prompt guide and improving your own prompts? Working with GLM but I think it's too wordy and adds too much the AI can't really interpret.
>>
>>109509949
Literally it. two videos into one. With audios. In standard h264 format Nothing else
>>
>>109509905
ur probably the dumbest poster itt, congrats
>>
>>109509905
>rely on AI
It's wasn't actually possible to create useful apps in 30 seconds before AI.
>>
He's the same bored retard trying to bait, you can tell by his constant adversarial tone. Ignore him or post wheelchair gens.
>>
>>109509905
>sigh.... you guys really rely on AI on literally everything
Slopping up a GUI wrapper over ffmpeg with your precise wants is like the most fit usecase for LLMs
>>
>>109509988
this. after the waifu app thing he pointed out it stopped being funny.
even if it were a real retard, it stopped being funny, just move on.
>>
>>109509905
says the guy who is too brown (low iq) to use fucking ffmpeg, lmao
>>
how car can you push the model off of the trained 15s limit theoretically? I want more
>>
File: ermmm achsually.jpg (38 KB, 736x611)
38 KB JPG
Is it worth getting anything above 32gb vram until you manage to cross the 256gb mark where you can run quant 4 versions of 500b models?
I just don't see the point of that dead man's land zone in between. MAYBE 48gb is justified as it usually comes in a standard card and allows you to run extra stuff ontop of your main LLM like additional small video gens or image gen.
>>
>>109510006
you can gen longer, its just too costly
>>
File: too strong for you.png (996 KB, 1040x917)
996 KB PNG
>>109510006
H3'S THEORATICAL LIMITS ARE TOO STRONG FOR YOU GOONER
>>
So Minimax H3 can't do explicit sex, but what about implied sex? Like sex under the covers, or with a POV camera? Or movie sex scenes? Have any of you guys tried it yet? Can you make it convincing?
>>
>>109509961
Alright. The AI is cooking
>>
>>109510014
>theoretically
I can trade off resolution or whatever, but how much longer are we talking before the gens fall apart?
>>
File: 00438-453709295610.mp4 (3.37 MB, 1088x1920)
3.37 MB
3.37 MB MP4
>>
>>109510038
Nice slowmo, homo.
>>
>>109510027
H3 can do anything if you reference the video and prompt properly.
>>
0.5MP, 20 step, 5s. Default comfy workflow: 384s. Adding patch sage attention, Mem Eff Sage attention, Low VRAM attention: 406s. T-thanks I guess?
>>
>>109509919
the third one...
>>
>>109510049
too much effort
it works, but we need a better approach. I think with some light loras and careful prompting we can achieve greatness
>>
>china releasing state of the art uncensored local models to destabilize the west
i love this psyop
>>
Do any of the turbo loras actually work with the reference model?
>>
>hurrrrr minimax can do anything with a reference!
then I may as well just watch the reference
>>
File: 1763926550209178.png (358 KB, 640x480)
358 KB PNG
>>109510071
>>
>>109510087
please be trolling
>>
r2v is fun

https://files.catbox.moe/0vl11z.mp4
>>
>>109510084
Yeah, all of them.
>>
>>109510084
yes >>109509614
>>
>>109510078
china won a long time ago. the playing field wasn't fair from the start. they arent heavily regulated, have no issue with straight up training their ai models off western sota models(deepseek/kimi), embrace ai as a culture and far stronger educational backgrounds. the west had no chance.
>>
File deleted.
>>
>>109510094
I'm very curious if you can get the lip smack from JC in there
>>
Alright that's it. I'm downloading more RAM.
swapoff /swapfile
rm /swapfile
btrfs filesystem mkswapfile --size 12G /swapfile
swapon /swapfile
>>
>>109510096
>>109510100
I will try, it seems to be more ridgid with references.
>>
I'm using a default workflow for minimax and my outputs have a lot of artifacting. Do i need to crank steps? Upping resolution doens't seem to really help.
>>
>>109510073
>too much effort
It's really not. Add a 5 second video. Say it's a weak_reference and it works as a lora for that concept.
>>
>>109510108
I think if you used that particular voice clip and prompted for it, it would probably work. my source is just a short .mp3 of random JC lines.
>>
>>109509652
>Anyone else tested it?
of course, it's the current best by far
>>
>>109510049
Wrong, idiot.
>>
>>109510121
yes
>>
>>109510128
Unlike you, I've actually been experimenting with R2V a lot, and neither reference nor weak reference works well for audio.
>>
File: 1755709497527592.png (190 KB, 277x337)
190 KB PNG
>python.exe
>>
least helpful general award goes to /ldg/
>>
>>109510159
Not your tech support Rajeesh
>>
punch this fuck
>>
>>109510159
u probably asked a brownoid question that an allm could have answered
>>
>>109510142
>>109510148
skill issue. Follow the prompt guide.
>>
>>109510171
this, bodied that saar freak
>>
>>109510159
lurk moar
>>
JC on global warming

https://files.catbox.moe/4zsxuz.mp4
>>
>>109510171
>>109510180
>>109510198
Dead internet theory proof #14
>>
>>109510183
I know more about the prompt guide than you do, brainlet. That's why I'm aware of it's limitations and issues.
Everything you say means nothing until you produce an example proving me wrong.
>>
>>109510215
yeah dead internet theory is when people don't spoon feed you basic questions that could just be binged.
>>
>>109510215
>being called out for being a retard
>must be bots
can't make this shit up.
>>
just asked my villages local facebook group and they said "pls redeem sir"
>>
File: 77777766666.mp4 (2.62 MB, 672x1216)
2.62 MB
2.62 MB MP4
>>
Did more testing with turbo and reference model, the model diminished in ability when working with more than one ref and more complex prompts
>>
>>109510215
u were too big of a pussy to link your question as you complain because u know it indeed was a brownoid tier question that an llm could have answered.
>>
>>109510240
impressive, now make them fart really loudly.
>>
>>109510244
You're not helping sperging like that, especially when a brown man opened local image gen models to begin with
>>
What's the reference limits? One picture on one video? 2 pictures on 1 video was scuffed. 4 pictures was scuffed.
>>
>>109510257
>no link
case in point, concession accepted.
>>
Ya I think I like this 2pass WF
https://h.uguu.se/odfOgEGh.mp4
>>
>>109510240
Kino
>>
>>109510244
>>109509740
>>
>>109510244
>>109510225
>>109510223
Dont talk to me, Clanker
>>
>>109510268
I'm not him based on your same combative nature you're playing both sides
>>109509988
>>
>>109510273
this was posted on preddit like a week ago or something
>>
File: 1804.png (219 KB, 568x633)
219 KB PNG
>>109509857
>lora mode needs to be set to stochastic
same shit
>>
>>109510102
I guess we'll see what happens. Historically China hasn't had much tolerance for companies who become so powerful they're like a branch of the government, certainly not if they're lead by headstrong people like Musk.
>>
>>109510280
No worries — I'll step back. If you want help with something later, just say the word.
>>
>>109510278
>whats 2 + 2, dont give me that 4 bullshit
u got ur answer, u just didnt like it
>>
>>109510293
i guess it's not built to work for h3, because it works for wan.
>>
>>109510285
was it? link? It's an old prompt but I don't post my shit on reddit.
>>
>>109510297
nooo saaar i need easy peesy app to plug in video so it come out like nolan film saaaaar please help
>>
File: screenshot.1786304472.jpg (190 KB, 1259x849)
190 KB JPG
>>109509961
Almost done. Excuse the AI for going overkill.
>>
>>109510280
Unfortunately I'm abliterated, so your anus is going to be obliterated.
>>
>see kino h3 gens that japs are posting on X
>they’re better than anything posted here
>>
>>109510010

Coming from 5090 + 5070 Ti. Yes, it is absolutely worth going for 48gb, especially if you only have a 5090 in your system.
Getting an additional 16gb card opens up the Gemma range completely along with other models in that region and will give you a shittton of more context.
It will also allow you to do stuff like have Gemma 12b running on the 16gb card to caption images while 5090 does the heavy lifting in video generation, something I noticed was pretty damn useful after playing with H3.
I'd say it's a pretty optimal combo for local.
Is it worth going higher than that by adding another 16gb? I doubt it, but then again I have no idea how much of a speedup you'll get in lower quants of Dsv4f by adding another card into the mix.
>>
>>109510325
AI slop, Japan: :)
>>
>>109510317
Thanks bro. Better than the rest of assholes here
>>
>>109510325
Japs are some of the biggest goyslaves, probably using maximum quality API H3 with promp enhancer.
>>
>>109510326
Too bad most cases can't fit a second card worth a shit, that's my biggest blocker from doing that
>>
>>109510317
which AI creates this exact style of user interface I've been seeing everywhere lately?
>>
>>109509679
???
Takes twice as long for s/it for me.
Likely won't wait to measure quality.
Not sure if a troll post or I am missing something.
>>
>>109510345
Most local models can do this, this is a simple wrapper and basic bitch UI
>>
>>109510156
>640GB ought to be enough to run any Python program
>>
>>109510351
yes nigge but what ai is it specifically that goes for that style
>>
File: screenshot.1786303239.jpg (138 KB, 803x476)
138 KB JPG
>>109510345
im using claude, but I just told it to make it modern. I didnt give it any other design tips
>>
>>109510340
if you download his vibe coded app you are 100% joining his proxy or botnet network.
>>
a bomb

https://files.catbox.moe/ptgekw.mp4
>>
>>109510355
Seeing how you're shitting up the thread non stop you probably can't afford to run it
>>109510361
I doubt it but it would be funny if his neg hole got pozzed
>>
>>109510377
Anyone can afford to runit
>>
File: 1755896194545504.png (37 KB, 2754x238)
37 KB PNG
does anyone have issues with git and github those last few days? it's unstable as fuck
https://news.ycombinator.com/item?id=49198302
>>
>>109510343

Yeah I know, most cases weren't built for these cards.
I have my 5070 Ti zip tied to the side of my case. It has the exact same cooler as the 5090 so there's no way in hell they fit in there, even though it's an ATX case, as they're basically 3.5 slot cards.
Only case that makes sense is basically the Phanteks Enthoo Pro 2 Server Edition.
That's guaranteed to fit the cards while leaving enough of a gap between them.
I'll buy one of these soon.
>>
>>109510326
I have 10GB 3060 still lying around before the upgrade to 5080 (the case is too small as the other anon mentioned)
does multigpu sampling works in comfy? I want better H3 gens. If yes, I'll buy a new case
>>
File: 724962893715490.mp4 (3.78 MB, 544x960)
3.78 MB
3.78 MB MP4
>>
>>109510325
I made lots of kino but janny wouldn't like it.
>>
>>109510361
nah, I just happen to be in a good mood thanks to a banger ref2va goon gen. it was delicious
>>
>>109510390
Have you tried codeberg or gitlab (local hosted)? Github is working ok for me and locally forego works rock solid
>>
File: H3i2v_00020_NoAudio.mp4 (1.73 MB, 672x928)
1.73 MB
1.73 MB MP4
>>109510362
Nice. Fun fact: twin towers are already missing from Deus Ex skybox despite being made before 9/11
>>
>>109510325
sasuga asian jeans desu
>>
https://files.catbox.moe/wfcu86.mp4
Two passes kinda work, but idk about samplers, sometimes shit deep fries too much.
>>
>>109510413
They just forgot to add them in?
>>
>>109510420
Hardware restraints.. or so MJ12 would like you to think
>>
>>109510417
https://h.uguu.se/rfGgBSte.mp4 (NSFW)
er_sde/beta57, 20pass but stop at 10
into turbo euler/beta start at step 4 out of 8

I was getting deepfried results when I let the first step go all the way.
>>
>>109510397

I tried multigpu but I couldn't get it working, threw a bunch of errors, so I have no idea how well that actually works.
Someone else can chime in there.
But when it comes to the LLM related stuff it's absolutely worth it stacking GPUs.
In general I would imagine we're moving towards parallelism eventually with all of the AI stuff, so it's only a matter of time until multigpu is a standard with all of these systems and keeping those old cards around will pay off.
>>
Use Picture 1 as the exact character identity reference for JC Denton, keeping his signature black sunglasses and stoic features. Extract the vocal timbre, frequency, and speech pattern from Audio 1 to build a voice clone.

Cinematic medium shot of JC Denton.

0 to 1 seconds: JC Denton is sitting in an art studio, in front of a blank white canvas and is completely silent.

2 to 5 seconds: JC Denton paints an anime style Hatsune Miku on the white canvas and is completely silent.

6 to 7 seconds: medium shot of JC Denton painting and is completely silent.

8 to 10 seconds: JC Denton looks at the painting and says "That's a nice Hatsune Miku, for sure.".

https://files.catbox.moe/jsd94a.mp4
>>
File: screenshot.1786305319.jpg (358 KB, 1272x809)
358 KB JPG
>>109510340
AI is complete. Feel free to try it.
This requires ffmpeg in path:
https://files.catbox.moe/nmojie.7z

This has ffmpeg bundled already:
https://files.catbox.moe/j922x9.7z
>>
>>109510447
his neovagina closed up
>>
>>109510447
eeeeew FUCK
>>
>>109510325
Not surprising. Japs are more creative on average, and creative people get the most out of AI.
>>
>>109510463
GPT-5.6 says this has like 5 backdoors in it
>>
>>109510463
I might test this later
Thank you for your contribution
>>
>>109510325
Examples?
>>
>>109510447
yeah, I'll try that, thanks
>>
>>109510463
Thanks for spoonfeeding me anon. I just want to make a mini scene of my successful gens
>>
>>109510463
>exe
>no source
yeah, anybody stupid enough to open that deserves to get cyber raped
>>
>>109510490
Post it or you're ungrateful fuck
>>
File: dtghnjdtghnjtdghnjtde.png (8 KB, 354x265)
8 KB PNG
>niggerlicious voodoo going on in comfyui just raped my gen speeds and i see why now
what in the
>>
>>109510498
Stop using bots
>>
honestly cant believe you niggers are genning on windows
>>
>>109510491
what's the point of including the src when the anon isn't a programmer?
>>
File: 1754827105726535.jpg (117 KB, 1000x915)
117 KB JPG
this shitpost was made by Debo. Know his post pattern
>>
https://d.uguu.se/UABTihpT.mp4
birthing machine
>>
File: file.png (10 KB, 468x97)
10 KB PNG
>>109510491
Yeah, where the fuck is the source code? 2 exes and one is just massive as fuck. Where is your source code? I won't be trying your program, as you compiled it for Windows and provided no source code repo I gave it the benefit of the doubt, and even downloaded both sip files, yet no source code. Crazy.
>>
>>109510305
>>109510293
Sorry, it works actually, thanks. I selected an fp8 model by mistake
>>
>>109510538
Why are you helping someone that begged like a bitch because he's too retarded to use a basic web search?
>>
>>109510538
Bro if you're going to whine for that, just ask chatgpt to make one for yourself
>>
>>109510550
Because I can and it's the right thing to do.
>>
I didn't see the namefaggotery, not replying further schizo
>>
>>109510554
You don't belong here, you're using a trip and feeding a troll that has nothing else going on and does this daily. You can identify him by his constant arguing and never having a solution to any subject
Read the OP
>>
Yeah, very happy with this workflow, time to gen some kinos
>>
>no gens
Your cums dried out along with your creativity
>>
File: 1768028321179400.mp4 (778 KB, 736x576)
778 KB
778 KB MP4
>>
>>109510538
one is massive because it has ffmpeg and other libs bundled in it. i didnt want anon complaining it didnt work because they didnt have ffmpeg in their PATH env variable.

the src is 2.5gb. just feels pointless posting it when i know you arent going to look at it
>>
>>109510574
you have no gens either thoughbeit
>>
>>109510491
Some source file being provided doesn't mean the compiled exe, which is what the brown anon who needed to be spoonfed is going to run anyway because if he had any brain he wouldn't be in this position in the first place, is in any way related to it. The way to go here is to decompile it and ask an LLM what that shit does.
>>
>>109510578
>the src is 2.5gb
HAHAHAHAHAHAHA
>>
>>109510578
>the src is 2.5gb
What the fuck. Is that just what Rust is like?
>>
>>109510594
if the nigga can't figure out how to vibe a simple app he isn't going to know how to decompile that shit nigga. this guy is 100% getting pwned one day.
>>
File: wallpaper_original.jpg (2.72 MB, 3378x1688)
2.72 MB JPG
>>109510594
Do you know about reproduciblity and the importance of reproducible builds?
>>
bollywood is finished.

https://files.catbox.moe/igw4b6.mp4
>>
>>109510615
pretty good
>>
>>109510615
>bollywood is finished.
More like bollywood is just ramping up
>>
>>109510604
whoops, I included multiple folders that had binaries in them. it's less than about 15MB.
>>
>>109510632
>it's less than about 15MB
hahahahahahaa. Still ridiculous, but slightly less so.
>>
File: screenshot.1786306929.jpg (68 KB, 802x213)
68 KB JPG
>>109510656
it's 99% cli.win32-x64-msvc.node from node_modules. needed for @tauri-apps/cli. either way im not posting sauce so either use it or dont
>>
Is a video or image better for copying a art style?
I want to copy how the anime looks
>>
>>109510578
>the src is 2.5gb.
ok now we have confirmed the retardation required to spoonfeed a brown to this level, u have to be even more retared urself
>>
If you're on at least 16gb, a 4000 to 5000 series, remove kijai's low vram node, 2ldr the math doesnt math up and causes inconsistent/slower speed, and remove all memory related startup flags, they're all memes that also slow it down and fuck with its memory management
2ldr for the 2ldr let cumfart manage its own memory, only use sage attention and the h3 sage attention patch. i just got my reference model speed at 0.8mp 10s down to 64s/it before the compile speed boost
thats.. basically every single speedup node removed for a speed BOOST. rough buddy.
>>
>>
File: 1623616564897.jpg (28 KB, 480x360)
28 KB JPG
this time, flux 3 lost. that stupid team made a mistake by ignoring us. local is now strong, even without flux 3
>>
>>109510517
Kino
>>
File: gnhbgxvhnbgxdh.png (41 KB, 981x294)
41 KB PNG
>>109510687
*actually 55s if i stop bringing up the 4chan tab kek
>>
>>109510686
>didn't read rest of reply chain
>u
>urself
>>
>>109510696
not only that, but their image model will probably have the same impact, I really can't wait
https://xcancel.com/MiniMax_AI/status/2086253065657790895?sort=Likes#r
>>
Whats the most complex h3 gen youve made so far?
>>
>>109510687
>>109510705
No one genning woth 16 gb of ram. Thats insane
>>
>>109510712
so Krea is dead long live MiniMax?
>>
>>109510712
Are you joking, holy shit. What a great year for this hobby
>>
>>109510720
I do. Works with Krea2 and H3.
>>
>>109510720
vram, hence 4000/5000 series you dingaling.
and there actually are a few here genning on 16gb of ram, with an 8gb vram card.
>>
>>109510719
a conga lone of women with dicks fucking each other in the ass
>>
>>109510740
share it bruv I wanna see
>>
>>109510732

no you are not fucking fag. why lie?
>>
JC drives a car and somehow breaks physics

https://files.catbox.moe/5fcxun.mp4
>>
>>109510751
Did you read the links in OP?
You're supposed to ignore his screeching, he spends all day baiting.
>>
>>109510687
>remove kijai's low vram node
Which one ?
>>
>>109510746
dont want jeets stealing my porn
>>
>>109510726
>What a great year for this hobby
I was about to call 2026 a dud but yeah, now it's a great year
>>
any llm can shit out a 5kb 120line python script to pass args to ffmpeg to both extract last frame of a video or combine two videos. can be written by literally just typing into google and bro waits an hour to download a 3gb rootkit instead of just learning how to pass a few args in his cli. fookin grim m8. has to be trollin..
>>
https://x.com/ArigatoLin557/status/2084993384306135094
>the superior japanese gens in question
>>
Someone prepare to make the thread
>>
File: ss_20260809_153029.png (136 KB, 1217x579)
136 KB PNG
>>109510751
???
>>
>>109510770
No shit do you even read OP?
It's the same pattern
>>
>>109510784
>>109510784
>>109510784
>>109510784
>>
>>109510757
>my vision is augmented.
idk why but this line cracks me up every time.
>>
>>109510757
>my vision is augmented
God damn it...
>>
>>109510732
>>109510781
They're bots ignore them.
>>
>>109510770
>anon already satisfied with his product
>sleek ui with plenty of options
>complaining about someone else's gift
yikes
>>
Guys can someone vibecode a multi GB interface for me as exe without sourcecode so i don't need to type in a ffmpeg command? I will run it as admin because i trust you. Thank you
>>
there we go, the car moves forward off the ramp.

https://files.catbox.moe/zkc0eu.mp4
>>
>>109510826
Working on it, sir.
>>
>>109510826
I'm assuming the requester and provider are the same person, in an attempt to get random anons to use the file

You need to be extremely lazy to ask other people make things with AI for you, when you can just ask AI directly instead
>>
>>109510837
Looks like a gmod animation
>>
>>109509770
Windows live movie maker
>>
My dumbass stuck scrounging for free credits and cheap GPU rentals because Nvidia told me 6 years ago 8gb of vram was enough! Somehow still 4x~ cheaper running serverless and 5-6x cheaper renting if you constantly Gen vs API costs at .17 cents/second.
>>
I been using image to video instead of references to video. I feel dumb.
>>
>>109511133
Yeah, reference is a bit slower, but way more flexible.
>>
>>109509939
Final Fantasy 8 too much
>>
>>109511133
>>109511139
what is the difference, surely image is the best for character consistency?
>>
>>109511892
Go to the new thread, this one is dead.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.