[a / b / c / d / e / f / g / gif / h / hr / k / m / o / p / s / t / u / v / vg / vm / vmg / vr / vrpg / vst / w / wg] [i / ic] [r9k / s4s / vip] [cm / hm / lgbt / y] [3 / aco / adv / an / bant / biz / cgl / ck / co / diy / fa / fit / gd / hc / his / int / jp / lit / mlp / mu / n / news / out / po / pol / pw / qst / sci / soc / sp / tg / toy / trv / tv / vp / vt / wsg / wsr / x / xs] [Settings] [Search] [Mobile] [Home]
Board
Settings Mobile Home
/g/ - Technology

Name
Options
Comment
Verification
4chan Pass users can bypass this verification. [Learn More] [Login]
File
  • Please read the Rules and FAQ before posting.
  • You may highlight syntax and preserve whitespace by using [code] tags.

08/21/20New boards added: /vrpg/, /vmg/, /vst/ and /vm/
05/04/17New trial board added: /bant/ - International/Random
10/04/16New board for 4chan Pass users: /vip/ - Very Important Posts
[Hide] [Show All]


Janitor acceptance emails will be sent out over the coming weeks. Make sure to check your spam folder!


[Advertise on 4chan]


File: 1770170970932083.webm (2.11 MB, 1143x2048)
2.11 MB
2.11 MB WEBM
Discussion and Development of Local Image, Video, and Music Models

Previous: >>109471487

https://rentry.org/ldg-lazy-getting-started-guide

>UI
ComfyUI: https://github.com/comfyanonymous/ComfyUI
SwarmUI: https://github.com/mcmonkeyprojects/SwarmUI
SDWebUI: https://rentry.org/ldg-lazy-getting-started-guide#the-stable-diffusion-web-ui-lineage
Wan2GP: https://github.com/deepbeepmeep/Wan2GP

>Checkpoints, LoRAs, & Upscalers
https://huggingface.co/models
https://huggingbay.xyz
https://civitai.com
https://civitaiarchive.com
https://openmodeldb.info

>Tuning
https://github.com/spacepxl/demystifying-sd-finetuning
https://github.com/ostris/ai-toolkit
https://github.com/Nerogar/OneTrainer
https://github.com/tdrussell/diffusion-pipe
https://github.com/kohya-ss/sd-scripts
https://github.com/kohya-ss/musubi-tuner

>Minimax H3
https://huggingface.co/Comfy-Org/MiniMax-H3

>Krea 2
https://huggingface.co/krea/Krea-2-Raw
https://huggingface.co/krea/Krea-2-Turbo
https://lumenastrum.github.io/clio-style-preview/gallery/

>Anima
https://huggingface.co/circlestone-labs/Anima
https://tagexplorer.github.io/
https://animadex.net

>Klein
https://huggingface.co/collections/black-forest-labs/flux2

>Misc
Local Model Meta: https://rentry.org/localmodelsmeta
Share Metadata: https://catbox.moe | https://litterbox.catbox.moe/
Txt2Img Plugin: https://github.com/Acly/krita-ai-diffusion
Archive: https://rentry.org/sdg-link
Collage: https://rentry.org/ldgcollage_v2

>Neighbors
>>>/aco/csdg
>>>/b/degen
>>>/gif/vdg
>>>/d/ddg
>>>/e/edg
>>>/h/hdg
>>>/trash/slop
>>>/vt/vtai
>>>/u/udg

>Local Text
>>>/g/lmg

>Maintain Thread Quality
https://rentry.org/debo
https://rentry.org/animanon
>>
https://files.catbox.moe/u6ou7n.mp4
we must be better men and not stack cache nodes!
>>
What's with all the troll bakes recently full of off-topic links?
>>
Reposting the question here, it's an interesting one
>>109472926
>Why is no one talking about how to improve the quality, only about how to gen faster?
>>
>>109472946
btw I haven't found anything better than res_multistep & simple, I asked Codex and it said that matches up fairly closely with what minimax recommends.
>>
>>109472944
why would the original OP be asleep at this time? does he live in india or something?
>>
Blessed thread of frenship
>>
I will NOT re-install cumfartui completely all over again just to run the latest model potentially faster.. No.. Please don't make me reinstall dependencies and redownload fucking torch again.
>>
>>109472972
Euler
>>
>>109472936
thanks for the bake
>>
>mfw Resource news

08/05/2026

>Inline Studio v1.2.62 - Minimax H3 Lora training still only
https://github.com/inlineresearch/Inline-Studio/releases/tag/v1.2.62

>Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-MiniMax-H3-GGUF

>Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF
https://huggingface.co/nif0/Qwen3-VL-32B-Instruct-ultra-uncensored-heretic-H3-GGUF

>MiniMax-H3-TAE: 2D tine VAE for MiniMax-H3
https://huggingface.co/Kijai/MiniMax-H3-TAE

>SPADE: An Input-Adaptive Sparse Attention Engine for Fast Video Diffusion Models Inference
https://github.com/6somehow/DAC-SPADE

>CAPE-T2V: Captioner-Anchored Prompt Enhancement toward Two-Sided Conditioning Alignment in Text-to-Video Generation
https://github.com/yizzz927/CAPE-T2V

>JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion
https://github.com/jd-opensource/JoyAI-Video-Edit

>ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs
https://github.com/YangYangGirl/ParVL

>OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
https://huggingface.co/JamesZar/OliveGemma-3B

08/04/2026

>stable-diffusion.cpp adds support for MiniMax-H3
https://github.com/leejet/stable-diffusion.cpp/blob/master/docs/minimax_h3.md

>ComfyUI Spectrum MiniMax H3: 34% lower Euler sampling time, 30% lower RES time
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

>MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
https://github.com/IntMeGroup/MIEScore

>PeCA: Palette Context Assisted Inference for Test-Time Paint-Bucket Colourisation on Animation Videos
https://rathgrith.github.io/PeCA

>Kandinsky WM 1.0: A family of models for Physical AI
https://github.com/kandinskylab/kandinsky-wm

08/03/2026

>MiniMax H3 Official Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)
https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md
>>
>>109472986
>re-install cumfartui completely all over again just to run the latest model potentially faster.
Source ? Is it really makes it faster ??
>>
>>109472965
you make it as fast as you can so you can iterate faster, the faster you iterate the better your workflow gets, the better your workflow gets the better your gens get, then you focus on quality.
>>
>mfw Research news

08/05/2026

>SeCo-SBIR: Semantically Consistent Prompt Learning for Zero-Shot Sketch-Based Image Retrieval
https://arxiv.org/abs/2608.03120

>HyperbolicDiffusion: Sharp & Scalable Tiled Generation on the Hyperbolic Plane
https://arxiv.org/abs/2608.03422

>DiverseDiT++: Quantifying, Analyzing, and Promoting Representation Diversity in Diffusion Transformers
https://arxiv.org/abs/2608.03082

>Can T2I Models Draw from the Right Frame of Reference?
https://arxiv.org/abs/2608.03357

>Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
https://tencent-hunyuan.github.io/Hunyuan3D-Buffalo1.0

>Self-Supervised Representation-Guided Generative Dataset Distillation
https://arxiv.org/abs/2608.03218

>Latent Reward Registers for Diffusion Preference Alignment
https://arxiv.org/abs/2608.03929

>UniWorld-Design: From Pixel Generation to Layer-Native Design
https://arxiv.org/abs/2608.03971

>MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding
https://arxiv.org/abs/2608.03708

>RIDGE: Re-Noising with Internal Dynamic Guidance for Image Editing
https://arxiv.org/abs/2608.03059

>Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending
https://arxiv.org/abs/2608.03269

>Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds
https://arxiv.org/abs/2608.03135

>TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models
https://arxiv.org/abs/2608.03057

>Adaptive Two-Stage Visual Token Pruning for Efficient Inference in VLMs
https://arxiv.org/abs/2608.03112

>Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
https://arxiv.org/abs/2608.03875

>Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
https://qwen-3d.github.io

>When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware
https://arxiv.org/abs/2608.03649
>>
>>109472953
loool, so which cache node is the best then? still spectrum?
>>
>>109472986
just pull it bro
>>
>>109472965
because if you can gen faster without altering the quality too much you can go for higher res with acceptable rendering times
>>
>>109473011
that anon's custom settings he posted for the cache node might put it slightly above spectrum
>>
>>109473022
the time is the same or is it faster?
>>
>>109472986
>reinstall
just update it dumbass
>>
sometimes Minimax makes the character talk gibberish just before it says its supposed sentence, how do you prevent that?
>>
redgjsieorjtg aijaw >>109473049 not sure what you mean.
>>
>>109473054
kek
>>
>>109473027
spectrum is definitely slower.
>>
>>109473007
this
diffusion generation is a spaghetti at the wall situation
>>
Any news about a 4 steps turbo lora? It's been one day already!!!
>>>/wsg/6208816
>>
>>109472989
that does seem better on a quick test, but I think I'd need to compare like 10 vs 10 generations to be sure.
>>
>>109473064
no way fag, spectrum was miles faster for me. but i THINK it's lower quality. i need to run one more gen (so it'll take like 5 minutes longer) to be sure though.
>>
>>109473075
>a 12.5gb lora
is he serious??
>>
>>109472972
>I haven't found anything better than res_multistep & simple
Same, I tried a lot of different settings but res_multi/simple always works.
>>
>>109473072
sacrilege!!
>>
the cache node might fuck with fine detail and prompt adherence
>>
>>109473072
Does it work the other way around, too?
>>
Can you handle my workflow anon?
https://files.catbox.moe/eb2xd3.mp4
>>
>>109473123
>might
figure it out
>>
File: Jen Wire.webm (3.92 MB, 704x1056)
3.92 MB
3.92 MB WEBM
>>109472180
Oh, you can make silent videos, but it's still gonna create an audio track in the output file, so you can't really turn it all the way off.
>>
File: 1782392482561125.png (174 KB, 363x240)
174 KB PNG
https://files.catbox.moe/b4b6ab.webm
>>>/wsg/6208834
>>
>>109473128
?potionseller refusing to sell his strongest potion
>>
tried to train innie lora for h3 results were kinda meh so didnt upload to civit but whatever here u go anons. dont @ me if it underperforms.
https://gofile.io/d/9fvLNJ
>>
File: 276172251.png (8 KB, 807x96)
8 KB PNG
what the hell is this nigga
>>
time to test out this newfangled "spectrum" you guys been talking about
>>
>>109473049
I think it has to do with the video being longer than however long it takes for the to say the specified dialogue. You either need to keep them talking long enough to match the vid length or clearly specify in the description what they are doing before or after speaking.
>>
Does anyone here have a workflow for using a reference image (a person) and copying the subject onto an existing video?
>>
>>109473170
someone is using your comfyui from outside your network. you know basic network security, right?
>>
>>109473175
test it out? You're already on it!
dohohoho!!
>>
Is there any upper limit for how long the reference audio for a voice should be?
>>
File: 00186-3311702336.png (605 KB, 640x512)
605 KB PNG
>>
File: MiniMax_H3_00360_.webm (1.8 MB, 896x704)
1.8 MB
1.8 MB WEBM
does /ldg/ think minmax is good with 3dcg?
>>
>>109473182
I was lazy and tried to just throw a 2 minute collection of a character's voice and it was really fucked up. Was a lot better after I took just one sentence.
>>
>>109473170
maybe this can help you, I don't have this anymore since I went for the --disable-api-nodes flag though
https://github.com/Comfy-Org/ComfyUI/issues/12618#issuecomment-3957464383
>>
File: yarra4.mp4 (1.46 MB, 416x736)
1.46 MB
1.46 MB MP4
>>109473189
sure?
>>
>>109473192
Oh damn. So just throwing more material at it to improve the quality isn't viable then.
>>
>>109473180
Why would a default install of comfy allow this to happen? No, I don't know network security, I am basically a consumer.
>>
>>109473181
`_`
>>
>>109473194
Thank you, I will try it out later.
>>
>>109473203
if you close your web browser / comfy tab while the web socket is open you're going to get that error
>>
>>109473203
uh oh, good luck
>>
lmao easycache raped the prompt adherence so hard, shit looked like a robotic powerpoint presentation
>>
>>109473198
tits not JIGGLY and BOUNCY enough
>>
>>109473218
Why would you use easycache over spectrum?
>>
The reward feels better because I waited so long to get a good local video model, hooray!!
https://files.catbox.moe/iz8b93.mp4
>>
>>109473227
trying both
easycache is kinda decent for previs to see if the prompt at the baseline does the required stuff
>>
>>109473189
https://litter.catbox.moe/jn103e.webm
>>
I really too indian and dont have creativity to do this.

How to stop being indians bros ? I wash my butt with water at least...
>>
>>109473229
GOD HE'S LITERALLY ME FR FR

>>109473159
man you just reminded me this character even exists, gonna make a vid of her shaking her ass or something while telling a joke.
>>
ok imma need the training config for that sweeny lora posted a few threads back, that shit is too good
>>
>>109473240
give an image to a LLM and ask it to give you 10 ideas, it can help
>>
still no sora though
>>
>8s 0,4mp ref2v
>10min
>0.5mp
>14m30s
god damn
>>
>>109473257
>ref2v
what all are you feeding into it?
>>
File: ComfyUI_2026-08-05_00040_.png (2.15 MB, 1024x1536)
2.15 MB PNG
>>
>>109473239
based, Daisy is the best waifu in Mario's land
>>
>>109473257
>>
>>109473265
8s video
>>
>>109473278
what resolution is the 8s video? because 8s ref on 8s gen is equivalent to 16s if the res is the same
>>
>>109473246
>sweeny lora
you can use a reference of her though, works well
>>
File: ComfyUI_2026-08-05_00044_.png (2.28 MB, 1024x1536)
2.28 MB PNG
>>
btw steps matter, 40 steps will take 2x long but if you're going for a best quality video you will see a big difference
need speedup loras!
>>
Is there an ideal image dimension for the reference images?
>>
>>109473284
yeah it's the same res
>>
for me it's 100 steps easycache
>>
>>109473298
obviously, we're coping with 20 steps and it's pretty good already
>>
>>109473298
ive been coping with 14 steps, is 20 steps a big step up?
>>
>>109473298
>>109473312
Why don't they release the turbo lora at launch?
>>
>>109473315
i haven't gone as low as 14 steps, you should just test it out. but doubling the steps you will see a big difference, i guarantee
>>
https://github.com/Comfy-Org/ComfyUI/pull/15334
>Adds suppport for int8_convrot quantized VAE for MiniMax-H3, tested that the VAE quality looks fine and it's about 1.5x faster.
meh, the vae decoding isn't that slow, I'll pass
>>
File: file.png (267 KB, 514x596)
267 KB PNG
>>109473300
one million pixels!
>>
>>109473315
>ive been coping with 14 steps, is 20 steps a big step up?
go for 20 steps + spectrum, you'll get something as fast as 14 steps but a better quality
https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
>>
>>109473334
i've already been using spectrum with 14 steps lol
>>
>>109473326
every second counts
>>
>>109473339
based.
>>
File: file.png (17 KB, 539x263)
17 KB PNG
>>109473326
shit go hard when this node light up u know u about to get lit
>>
>>109473298
>btw steps matter
Jesus someone rape these newfags already
>>
>>109473356
anon... it's the newfags i'm trying to educate. if you haven't noticed we have an influx of like a billion of them every new model release
>>
>>109473334
okay ill try out the spectrum meme. what are the best settings?
>>
https://files.catbox.moe/e14qfm.webm
>>
>>109473375
the ones you get when you load the node
>>
>>
>1 hour wasted on failed attempts
kino
>>
>>109473387
thats why we need the turbo lora so we can iterate faster
>>
https://files.catbox.moe/9jq2rj.mp4
>>
isn't spectrum capped at 20 steps?
>>
>>109473409
no you can go for any steps that you want, but going for too low won't really work, spectrum will consider that all the steps are important
>>
https://files.catbox.moe/syp9o3.webm
>>
>>109473386
lol wish /adt/ had more of this type shit
>>
https://files.catbox.moe/0q609f.mp4
>>
>>109473370
Can I still rape you tho
>>
>>109473448
>>109473424
he just likes the smell of an unwashed bellybutton
>>
>>109473454
NTA but you have my permission.
>>
>>109473470
can't be rape if there's a permission though
>>
https://files.catbox.moe/wuis74.mp4
>>
>>109472999
>>109473010
thanks!
>>
>>109473179
it's a shot in the dark right now. And I think it depends more on prompting than workflow.
>>
>>109473479
I take it back then. You do NOT have permission.
>>
>>109473480
Now this is what a peak gen looks like. Everyone else can just delete H3 now.
>>
H3 actually doesn't output exact duration. Sometimes you get trailing 0.35 seconds or something and it can fuck your alignment up if you chain multiple together. FYI
>>
>>109473507
the reason
>max(5, round(a * b)) + (5 - (max(5, round(a * b)) % 17)) % 17
>>
would be a great advertisment video for Minimax lol
https://files.catbox.moe/18inkw.mp4
>>
>>109473507
it's 17k+5, the math is in the note connected to duration
>>
are abliterated text encoders just snakeoil slop?
>>
I don't know how to run MiMax on 12gb vram and at this point I'm too afraid to ask
>>
>>109473524
yes, it's not censored to begin with, the heretic guy just blanket applies uncensored slopping to every LLM out there. it just makes it stupider.
>>
File: MiniMax_H3_00005_.mp4 (593 KB, 480x864)
593 KB
593 KB MP4
It can't deal with reflections properly, or I cannot make the prompt clear enough make it work.
>>
Having a hard time trying to prompt barefoot footsteps sound. H3 always wants to make it sound like high heels
>>
>>109473529
you download the int8 model and you let comfyui do the automatic RAM offloading for you, if you have OOM, try those flags
>--vram-headroom 1
>--disable-pinned-memory
>>
>>109473494
https://files.catbox.moe/h77iw4.mp4
>>
>>109473532
it definitely feels like some things were omitted from the text prompts in the dataset, but no videos outright culled from it either, leading to some concepts just... happening when the model feels like it
>>
>>109473535
plod, thump, uh.. thud
>>
>>109473534
Prompt issue
>>
>>109473170
what are you running it on, are you doing it remotely? looks to me like you just got owned. You should be connecting to comfyui remotely as if it was a local IP through SSH or something and not just opening a port. Now the question is how fucked are you?
>>
>>109473535
try more descriptive, like "to the soft padding of feet on carpet"
>>
File: settings.jpg (23 KB, 441x400)
23 KB JPG
>>109472883
alright, quick update
Copied these settings from [spoiler]reddit[/spoiler] and now I'm getting much faster generation time. Did some comparisons and the quality hit was neglibile, almost like a variation of the same thing.

No H3 Cache, 0.3MP @ 16:9, 10 seconds: 647s
H3 Cache Settings from last thread: ~450s
H3 Cache settings fro Reddit: 372s
>>
>>109473203
are you attempting to run comfyui through 2 different web browser tabs?
>>
>>109473534
the gunt is fucking disgusting
>>
>>109473553
if you pulled master these are now the default settings for the node.
>>
it can do ASMR too
https://files.catbox.moe/st0f5h.mp4
>>
>>109473542
What's more likely is that the LLM-hallucinated prompts aren't that accurate
>>
>>109473561
can it do duke nukem asmr
>>
>>109473561
what kind of prompt for that? i like it
>>
File: 1770090312473959.webm (2.34 MB, 896x1920)
2.34 MB
2.34 MB WEBM
no audio so it doesnt fry your brain
>>
>>109473574
0:00–0:02
Mickey Mouse sits at a tiny ASMR desk in a cozy, softly lit studio, positioned extremely close to a binaural microphone. He gently taps the microphone with his white-gloved fingertips and whispers in his unmistakable high-pitched Mickey Mouse voice:

“Gosh… you made this?”

0:02–0:05
He slowly leans closer, maintaining cheerful eye contact while softly crinkling a sheet of legal paper beside the microphone:

“Well, I’m gonna have to sue ya…”

0:05–0:08
Mickey pauses, gives the camera a playfully threatening smile, then whispers:

“Ha-ha! See you in court!”

He finishes with his distinctive “Ha-ha!” Mickey laugh, followed by two delicate microphone taps.

Exact classic Mickey Mouse appearance, voice, mannerisms, white gloves, red shorts and yellow shoes. Intimate stereo ASMR audio, soft whispering, gentle paper sounds, detailed glove tapping, no background music.
>>
>>109473579
catbox? curious what the prompt was
>>
>>109473561
AAAAAH DISNEY SAVE ME
>>
File: best model ever!.png (123 KB, 555x360)
123 KB PNG
https://files.catbox.moe/65o65y.mp4
>>
>using params from plebbit
>thinking spoilers work on /g/

newfag holocaust when
>>
>>109473583
Then ask for prompt and not catbox.
>>
>>109473590
i do not like the way that miku looks
>>
>>109473581
thank you
>>
>>109473594
walls of text are unsightly
>>
File: 1776043020789323.jpg (149 KB, 830x735)
149 KB JPG
>>109473583
>>
>>109473604
Damn H3 threw in the jiggly boobies for FREE?
>>
i... last one, i promise
https://files.catbox.moe/nsxltn.mp4
>>
File: h3 be like.png (562 KB, 1280x720)
562 KB PNG
>>109473608
>>
>>109473559
ah... fml
>>
>>109473610
kek
>>
>>109473579
>no audio
I want the audio.
>>
>>109473610
hi Lodestone!
>>109473601
I had a better migu on the first try but she appeared before he asked for it, I don't want to wait 10 more mn, take it or leave it :(
>>
>>109473591
>NOOO YOUR SUPPOSED TO SPEND ALL YOUR TIME ON ONE BOARD
a bloo bloo bloo gramps, a bloo bloo bloo
>>
used to be that redditors would steal content from 4chan. now we steal content from reddit. Grim.
>>
what settings should I tweak on spectrum to match the speedup done by easycache's default settings? i'm kinda clueless what any of them do
>>
File: a6-06.jpg (22 KB, 1200x600)
22 KB JPG
>skip steps: 100%
>>
>>109473610
cute dog
>>
>>109473632
and discord, are anons unable to make kinos anymore?
>>
I have no idea what im doing with my prompts. Granted im using text-only so this is probably not as trivial, but honestly even with the guide for prompting they supply I feel fucking retarded.
https://files.catbox.moe/ia1b4n.mp4

Is there any particular short tip that you have?
>>
>>109473635
>looks at blank result
>"ah, it's exactly what I envisioned in my head"
>>
>>109473632
>reddit
everything is reposts from trooncord
>>
File: H3_Trailing_frames.png (53 KB, 834x1204)
53 KB PNG
>>109473507
FYI for multiple shot chads out there.
>>
trying to get h3 prompting down, there are two guides in the docs. should I be using one over the other? I imagine one is for exact super specific prompts and the other for more general vague prompts ?
>>
>>109473645
one is for the FL2VA model
one is for the REF2VA model
>>
>>109473632
>>109473643
almost as if ledditors and trooncord users are also using 4chan and are crossposting their videos or something...
>>
>On Monday, Tokyo police announced the arrest of a 32-year-old man from Himeji, Hyogo for using generative AI to create nude images of female track and field athletes — some of them as young as 12 — and posting them to social media groups. “It made me happy that I could make people happy,” the suspect reportedly told police.
>Investigators say he created around 300 images and posted roughly a third of them, some at the request of a 17-year-old member of the same group, who has also been referred to prosecutors. One victim is now in her 40s; the source photo used to create her deepfake was taken when she was in her first year of middle school.
>Japan has seen a string of deepfake-related arrests over the past year, including a man accused of selling more than 500,000 AI-generated sexual images of celebrities and another accused of creating deepfakes of more than 260 female celebrities. Each case has relied on a different legal theory, exposing the same underlying reality: while generative AI has transformed how these images are made, Japan’s legal response is still built from laws written for something else.
https://www.tokyoweekender.com/japan-life/news-and-opinion/man-arrested-over-ai-generated-nude-images/
>>
>>109473648
ahh that makes sense, ty anon I was very confused
>>
>>109473651
>500,000 AI-generated
jesus this dude never sleeps or what?
>>
>>109473651
>and posted roughly a third of them
well, duh
>>
>>109473651
just don't post your shit online, it's that simple
>>
>>109473651
>sharing the gens online
>using real people
>>
>>109473651
whats the moral of the story?
>>
>>109473651

>Has go into games machine
>Proceeds to share his waifu with others like a cuck in those hentai
>get arrested.

Deserves the rope.
>>
>>109473651
that's quite a collection
>>
anyone been playing around with steps in minimax? 20's the default but how low can you go with negilible losses?
>>
File: file.png (12 KB, 260x109)
12 KB PNG
Do seeds affect gen time?
Same settings prompt but a +1 difference in seed
>>
>>109473651
>It made me happy that I could make people happy
He didn't deserve it :(
>>
>>109473651
>>>109473653 >>109473655 >>109473657 >>109473658 >>109473660 >>109473665 >>109473670
least obvious botted replies
>>
>>109473676
beep
>>
>>109473676
What would be the purpose of the bot?
>>
File: 1782578419481702.png (126 KB, 498x487)
126 KB PNG
Why genning with Minimax cause my GPU fan to load 90% while genning with LTX only like 50% ?? I thought it only uses VRAM ????
>>
>>109473534
>1slag
>>
>>109473651
based
/ldg/'s ultimate boss
>>
>>109473676
>anon posts headline about ai crime
>anons respond to it cuz the issue here was obvious
>Umm.... le botting?!?!
>>
why does non_diegetic_music pause when ever people talk, how to stop that shit because its really fucking annoying and stupid. i'm not fond of this boomer prompting since i've had better results just prompting it normally except when the subjects talk, then its better to do something like.
the man says: (S1) <d> [English] text. But then again I've had some success with just literally. The man says "text"

Audio: sound happens

also seems to work so what the hell gives? Its not at all easy if it doesn't even work as described in the documentation. I'm gonna try some different methods, perhaps i can be structured with action in the video as time stamps and then followed by new paragraph with time stamps for dialect followed be another section with timestamps for the audio.
>>
>>109473684
>he doesn't pin his fans to 100% when generating
enjoy your broken gpu
>>
>>109473676
you must have autism or something lmao
>>
>>109473651
Nah. The Tokyo police are pervs too. They just want an excuse to talk to the purported "female victims".
>>
>>109473672
I'm using 15 steps just to save some time, 10 is too low aside from a rough test of a prompt. And it's a 480p as well, only on a 5060ti 16gb so not blessed with the speed of the better cards. I mean really 20 steps should be used. But this early on with still trying out prompts etc and needing a few gens done per prompt I take the hit on quality. To me it's about the content generated not just the visual quality anyway
>>
>>109473645
>>109473648
One is base that works for both, for things like camera control.
>>
Is the anon who made the rimming videos earlier still here?
>>
File: h3_00005.mp4 (3.4 MB, 640x1152)
3.4 MB
3.4 MB MP4
>>
>>109473695
probably try something like 'music plays uninterrupted as'. it'll halt action states at shot transitions, too.
>>
>>109473551
I am just using the .bat from the portable version. What's the worst that can happen to me now?
>>
>>109473653
Easy to automate a SD.1.5 script on Comfy or Auto1111 on a good GPU and make thousands of images per day.
>>
https://github.com/SandAI-org/MAGI-2-preview
https://sand.ai/blog/magi-2-preview
>114B total parameters
>no video showcase
lmaoooooo
>>
>>109473735
look at the requirements
>>
>>109473674
.
>>
>>109473710
you're here, so yes
>>
>>109473742
No, I had a question for them
>>
>>109473674
if you use the cache nodes (not spectrum) it's possible one seed skips less steps.
>>
Sigma chads, what settings are you using, and are you getting a noticeable improvement to audio or visual quality
>>
>>109473743
rimjobanon here. what's your question?
>>
>>109473759
8/4, yes.
>>
>>109473735

Wut? What a disgrace to use that venerable name with no result to show for it.
>>
>>109473284
So I can lower my ref video res and speed things up?
>>
https://files.catbox.moe/w8cvp7.mp4
how do I STOP it adding sound effects?
I told it:
overall_soundscape:
No ambient, environmental, or physical action sound is present at any point; the wind, footsteps, impacts, and explosion play silently. The theme song is the only audio in the entire video.
but it still added punches...
>>
>>109473743
>them
>>
>>109473735
>we've made a 100b parameter model
>that activates 6b per/tk
>and it needs 8 HOPPER gpus https://www.techpowerup.com/gpu-specs/nvidia-gh100.g1011
>all of them
wat
>>
if i am not running blackwell, should i still use nvfp4 or should i just use q4?
>>
>>109473266
Why is every model so bad at placing the smoke for cigarettes?
>>
>>109473772
don't you know the first rule of prompting is to avoid saying shit like
>Don't do X
>>
>>109473735
>>109473775
Imagine where we would be if AMD was on the table for this shit
>>
>>109473772
Sound prompting seems to be pretty buck wild. I have gotten it to add sounds only inconsistently despite matching the prompting guide. It almost never misses dialogue but sound effects just don't show up half the time, or more. You could try "n/a" in the overall soundscape field but it seems to be a bit voodoo.
>>
>>109473777
sometimes smoke can travel up the cig like that
>>
>>109473775
How does a 100B model require 8 80GB VRAM GPUs?
>>
>>109473513
how long did it take you to make that? What was the process?
>>
>>109473784
>if AMD was on the table
AMD is a controlled opposition
>>
>>109473735
114B is a small model in /lmg/
>>
>>109473763
I was wondering how you inserted a video into the R2V node, but it seems like you used comfyhelpersuite after searching a few threads earlier.
>>
>>109473579
is this overwatch?
>>
>>109473818
fortnite... overwatch... same shit
>>
>>109473558
me gusta
>>
>>109473651
>posting them to social media groups
dude fell for the modern day honeypots
I avoid those places like the plague
>>
>>109473806
NTA but in default comfyui there's a load video node and another node for splitting the video into image and audio. Use both and then route them into the respective slots.
>>
What node do you use to load a video in r2v? the comfycore load video node has a grey output but R2v takes a blue input.
>>
Please chinkmoot

Banish these newfags to the shadow realm


Thanks,
Anon
>>
/ldg/ got chinese cultured
>>
>>109473842
Video Helper Suite, custom node.
>>
>>109473849
Huh? Isn't that Kijai's node pack? Why would Minimax R2v require this?
>>
>>109473759
pixart sigma?
>>
>>109473853
good question
>>
So I heard h3 makers were going after anyone making nsfw and generally banning that usage, which means places like civitai will not allow it either.
I guess that means the model will be doa long term for "complete" nsfw?
>>
>>109473853
it doesn't. You can use Get Video Components. But the custom node is better.
>>
so then, after a few days of everyone niggerrigging and niggergenning, which scales hardest against our s/it? runtime or resolution? getting hit harder trying to gen at 1mp less than 10s, or 10+s at 0.5mp? where's the middleground?
>>
>>109473771
you can lower res, lower fps too if you want
>>
>>109473864
looks like we're fine, there's already a lot of nsfw h3 loras on civitai and none of them are nuked
>>
>>109473842
Load Video -> Get Video Components
>>
>>109473803
that's captialism for ya
>>
>>109473881
Oh nice, then what happened, last I read they were on a nsfw hunt?
>>
>>109473853
gee, anon, I don't know. Maybe the colors and input names mean something.
>>
>>109473892
>>109471258
>>
File: 1767604619045190.mp4 (873 KB, 864x480)
873 KB
873 KB MP4
>>
>>109473892
apparently civitai got a loicense from em
>>
>>109473892
only on hugging face so far.
>>
>>109473899
>tfw you remember you have free will
>>
>>109473884
TY anon this is what I needed :D
>>
File: 1771913871013376.jpg (16 KB, 519x299)
16 KB JPG
kino incoming
>>
File: 1770276765617761.png (3.6 MB, 1685x2048)
3.6 MB PNG
>>109473909
>you have free will
you don't
>>
>>109473553
>>109473559
fuck. Alright, fine, I'll add it to my testing. Can I add it through the manager, or do I have to git clone?
>>
File: MiniMax_H3_00024.webm (792 KB, 832x480)
792 KB
792 KB WEBM
If you want quality while saving vram, use the gguf quants of H3.
The Int4 variants gave garbage outputs (like really bad, as if the conversion didn't work correctly). Int8 may also have this issue.

The gguf quants are slower, but I actually get decent outputs. After applying sageattention 2.2.0 and Spectrum with conservative settings, I cut gen times down from 15~mins to 8~mins on a rtx3060 12gb and 24gb of system memory for a 5s I2VA clip at 0.4mp.

Using H3 Q4 quant and nif0's qwen3vl 32b ultra herectic IQ2_XSS with his forked gguf loader that adds support for the mmproj vision file. Also using molbal's gguf loader for the main model so I had to rename nif0's gguf loader custom node folder to tell the difference.

Standard UNet GGUF loader wont work here as it's using nif0's node, using the dynamic vram version using molbal's node instead.

https://files.catbox.moe/uxb7x1.webm
>>
>>109473929
>5s
>480p
>ggufs
nigga please fuck you talm 'bout
>>
>>109473791
n/a didn't work either. Maybe it's just not built for non_diegetic sounds only.
https://files.catbox.moe/gk65xc.mp4
>>
File: 1754650565787447.mp4 (742 KB, 864x480)
742 KB
742 KB MP4
>>
>>109473929
Weren't you that fag claiming LTX still has a usecase or am I mixing up my faggotposters
>>
>>109473929
mmm no
>>
>>109473804
is not like we can use multiple gpus in ldg
>>
File: 1770290282312856.webm (3.86 MB, 1920x912)
3.86 MB
3.86 MB WEBM
Holy shit a turbo lora is already here
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora

here's a lora strength comparison, it's 8 steps + euler + simple for all
>>
33 mins for a 20s vid fuck outta here lol
>>
>>109473951
fuck you it takes me that long for an 8s one
>>
File: p56.jpg (86 KB, 405x720)
86 KB JPG
>>109473950
based.... so based...
>>
>>109473929
can you post an example that has more movement than a flapping mouth and slight arm move
>>
>>109473950
Still a bit undercooked? The hands...
>>
Ok, the leddit settings for the h3 cache node definitely help gen times for me. 38% faster in one trial. There might be quality loss, but it's hard to tell - as in, it's not noticeable compared to ordinary seed variation.

Comparing:
INT8 convrot -> minimax sage attn patch -> spectrum
INT8 convrot -> minimax sage attn patch -> cache node

.3mp square, 15 seconds

Cache: 350 seconds
Spectrum: 482 seconds
>>
>>109473950
wtf? how does it look better than no lora?
>>
>>109473975
>how does it look better than no lora?
because it's at 8 steps for all anon, regular minimax at 8 steps looks like shit
>>
bros I didn't even ask her to lick her finger...she just has a mind of her own!

https://files.catbox.moe/2mj4wt.mp4
>>
>>109473772
>overall_soundscape:
overall_soundscape: N/A
>>
I have another R2V question.
How do you reference the audio from a video reference?

The prompt guide says:

>2.5 Visual and Audio Tracks from the Same Reference Video

> <Video N> and <Audio N> are numbered independently. Each index indicates only the label's order within its own category and does not encode a pairing between the two categories. The same reference video may therefore correspond to <Video 1> and <Audio 2>; different indices do not prevent them from coming from the same source asset.

>An ordinary reference video does not create <Audio N> merely because the file contains sound.

>An <Audio N> definition primarily states the audio's role and does not have to name the <Video N> it comes from. State the shared source only when needed to remove provenance ambiguity, for example:

> <Video 1> is the source video for the target video edit.
> <Audio 2> is the synchronized audio track of <Video 1> and is reused in the target video.

But as seen from the node (pic related) there is an input for "ref_video_audio_0" which increments upward.
If I reference <Audio 1> with pic related, will that reference the "ref_video_audio_0" input, or the empty "audio_0" input? The documentation doesn't make this 100% clear.
>>
>>109473950
>>109473975
i still dont understand how turbo loras work. why dont the model creators just make it turbo from the beginning, is that what zit did?
>>
>>109473988
>why dont the model creators just make it turbo from the beginning, is that what zit did?
yeah and it makes the model impossible to train
>>
>>109473988
I didn't for a long time and I was very skeptical, sounded way too good to be true
>>
>>109473898
>>109473900
>>109473903
Thanks anons, well good thing then.
Big websites will get their licenses and people will do whatever they want.
>>
>>109473950
Perhaps my 10 second gen dreams will come true after 3 days(?) Shit is moving fast
>>
wtf, to make it loop you just have to tell it to loop.

>Preserve a seamless loop and stable camera.
>>
File: 07591726.jpg (716 KB, 1792x2400)
716 KB JPG
>>
>>109473772
>The theme song is the only audio in the entire video.
use
non_diegetic_music: None interrupted fast paced techno drum beat with bassline.

Or what ever.
>>
>>109473985
Prompting starts from 1. The node slots is just coded to labeled from 0.
>>
>>109473988
Better to have the makers focus on the full fat thing then everyone can do whatever optimizations and quants they want after
>>
>>
>>109473985
from my understanding they are all in the "<Audio>" queue but all the ref_videos are added to the queue first. so the ref_video_audio_0 is Audio 1 and ref_audio_0 is <Audio 2> if both connected. If no video ref_audio_0 is <Audio 1>
>>
>>109473950
https://huggingface.co/larryvrh/MiniMax-H3-Turbo-Lora/discussions/1#6a73cf519b0aee71a9a71bcd
>Nah, this is just an experimental demo, not comyfui compatible yet. Will add comfyui support once the full training finished.
motherfucker I wanted to try it...
>>
>[WARNING] lora key not loaded: token_refiner.blocks.1.attn.qkv_proj.lora_A.weight
>>
ok so turbo lora at 4steps euler/simple
gibberish
>>
>>109474019
it's in the oven, people are working on it, relax
>>
>>109474019
>not comyfui compatible yet.
bruh.
>>
File: 1763664806370305.mp4 (763 KB, 864x480)
763 KB
763 KB MP4
>>109473950
>>
im going back to ltx. h3 is not mature yet
>>
>>109474028
>it's in the oven
he's posting it on huggingface though, so he wanted us to see it and try it
>>
>>109474038
see >>109474019
explicitly says it's happening but it's not ready
from this we learn that it's happening, see
>>
>>109474032
lmao why the fuck did the baby getting thrown 3 seconds in make me laugh harder than your actual original gen
>>
>>109474037
>im going back to ltx. h3 is not mature yet...you sick fucks!
>>
>>109474032
holy sovl
>>
>>109474032
holy fucking white teeth
>>
>>109474037
He says while everyone is posting kinos left and right that not even the top cloud models could produce.
>>
>>109474032
>i'll give you a 5090 if you spin around and throw your bab-
>>
>>109473950
wooooooooooow
>>109474032
kekk. more lol
>>
>>109474054
you're inviting a british joke
oh hell I'll do it myself; >t. british
>>
>>109474032
that's how the game should have been desu, imagine raising someone else's kid, yikes
>>
>>109474032
>When you don't get child support for McDonalds
>>
>>109474071
kek to the extreme
>>
>>109474019
someone get claude to convert that generate.py into a comfy node.
>>
what are you genning with H3 RIGHT NOW?
>>
>>109474032
>she spinned so fast she duplicated him
kek, the laws of physics are weird when we're getting close to the speed of light I see
>>
>accidentally wrote fish instead of fist
oh man
>>
In case flux3 isn't completely shit, can you technically use the output of minimax to teach it concepts?
>>
>>109474084
tacky mind control fetish nonsense
no other model has ever known how to do this out of the box even if it's all b-movie tier acting
>>
>>109474084
more migu >>>/wsg/6208856
>>
>>109474019
i used it just now but the quality sucks, guess that's what they meant by comfy support?
>>
File: 57543.png (142 KB, 1844x630)
142 KB PNG
>>109474084
kinos
>>
>>109474103
>i used it just now
you can't, it's not comfyui compatible
>>
>>109474084
Alien videos that could fool boomers
>>
>>109474107
your mom's not comfyui compatible but yknow what i made it work.
>>
>>109474113
>>109474113
>>
>>109474099
a man of culture >>6208208
>>
File: hot_girl.mp4 (711 KB, 576x896)
711 KB
711 KB MP4
>>109472741
>>
>>109474084
black_widow_hulk.gif
>>
File: 1777365224645269.png (347 KB, 750x310)
347 KB PNG
>>109474112
>>
>>109474118
Kek
>>
>>109474118
kek
>>
File: whats this..jpg (30 KB, 657x362)
30 KB JPG
>>109474107
?
>>
>>109474118
that's me
>>
>>109474133
now try the same gen without the lora, same steps and seed.
>>
>>109474133
the format of the lora is not comfyui compatible, look at your cmd console you'll get layers because the names of the layers aren't the same of what comfyui is expecting, this lora needs to be converted
>>
>>109474138
I did, and it worked but the quality was ass
can't post cuz it's pron
>>
>>109474084
tailjob
>>
>>109473988
I don't really know how it works but i imagine its distilling the model using many small clips of many concepts and if they concepts happen to be in your prompt the model can then shortcut based on what is inside the lora. Someone correct me if I'm wrong. The side effect is the model have less flexibility.



[Advertise on 4chan]

Delete Post: [File Only] Style:
[Disable Mobile View / Use Desktop Site]

[Enable Mobile View / Use Mobile Site]

All trademarks and copyrights on this page are owned by their respective parties. Images uploaded are the responsibility of the Poster. Comments are owned by the Poster.